AI & Automation · Houston, TX

AI that reads the paperwork so your team doesn't have to.

Every operation we walk into has the same job hiding in it: a person reading a document and typing what it says into a system. Invoices, BOLs, POs, emails, voicemails, photos from the field. That job is now automatable — reliably, affordably, and if it has to be, entirely on your own hardware. That's the AI we build for Houston businesses. The rest is mostly demos. (When the paperwork in question is inbound leads and follow-ups, that's our customer intake & outreach practice.)

Sound familiar?

Eight symptoms, and what they actually mean

None of these need a moonshot. Most of them need a model, a review queue, and some plumbing.

Drowning in manual data entry

Somebody's whole job is reading a document and typing it into a system. That job can now be checked by a person instead of done by one.

Documents get typed twice

The PO arrives as a PDF, gets keyed into the ERP, then keyed again into a portal. Every retype is another chance to be wrong.

Tribal knowledge trapped in inboxes

The answer exists — in a 2019 email thread, a shared drive, or the head of the one person who has been here twenty years. Nobody can search any of it.

You can't send data to the cloud

Contracts, compliance, or common sense say your documents stay in-house. That rules out the SaaS tools — not AI itself.

You tried ChatGPT. Now what?

Copy-paste into a chat window proved the concept. Wiring it into your real systems, with controls and an audit trail, is the part that pays.

Everyone wants an AI strategy

The board wants an answer and vendors want a signature — before anyone has checked whether the data can support any of it. Start there instead.

Nobody answers after five

Calls and voicemails stack up overnight and get triaged at eight — by which time the customer has already called somebody else.

The job comes back as photos

Site photos, a handwritten ticket, and a text thread. Someone at the office turns that into the record days later, from memory.

Capabilities

Practical AI, wired into real systems

Ten things we build, from document extraction to phone calls to forecasting. Every one ships with validation rules, a human veto where it matters, and documentation your team can run without us.

Document & Email Automation

LLMs reading the paperwork your team currently retypes.

  • Extraction from invoices, BOLs, POs, rate confirmations, and packing slips
  • Email triage: classify, route, and draft replies for a human to approve
  • Structured output validated against your business rules before it touches a system
  • Confidence thresholds — the ambiguous ones go to a person, not into the database
  • Works on PDFs, scans, and the odd formats your partners actually send

Private & On-Prem LLM Deployments

Self-hosted open-weight models for data that can't leave the building.

  • Llama- and Qwen-class open-weight models running on your hardware via Ollama
  • No per-token fees, no data leaving your network, no vendor reading your documents
  • Sized honestly — a workstation-class GPU covers more than most vendors admit
  • Same automation patterns as the cloud APIs, minus the compliance conversation
  • We run local models in our own operation — this is practice, not a brochure page

AI Inside Your Existing Systems

Copilots and answers wired into the ERP, WMS, and CRM you already run.

  • Retrieval-augmented generation over your SOPs, contracts, and manuals — with citations
  • Copilots in Power Platform, Teams, or wherever the work already happens
  • API-level integration with ERP, WMS, TMS, and CRM systems
  • Vector search over company documents that respects who is allowed to see what
  • Legacy systems included — if it has a screen or a database, we can reach it

Process Automation with AI in the Loop

Automation where the model does the reading and a human keeps the veto.

  • Triage queues: AI sorts, drafts, and flags; people approve
  • QA gates that catch the model's mistakes before your customers do
  • Human-review workflows with an audit trail of who approved what
  • Escalation rules for the cases the model should not decide
  • Accuracy measured over time, so trust is earned rather than assumed

AI Agents & Multi-Step Workflow Automation

Software that carries a whole task through, not one prompt at a time.

  • An agent that reads the email, finds the order, checks the portal, and prepares the update
  • Tool access scoped deliberately — it reaches what the job needs and nothing else
  • Dry-run first: it proposes every write for a week before it is allowed to make one
  • The deterministic steps stay deterministic; the model only handles the judgment calls
  • Every action logged against the input that caused it, so a bad run can be traced and undone

Voice, Phone & Meeting AI

The calls, voicemails, and meetings nobody has time to write up.

  • After-hours answering that captures the details and files a real ticket, not a callback slip
  • Voicemail and call recordings transcribed, summarized, and attached to the right customer
  • Dispatch and service calls turned into work orders with the address, the times, and the parts
  • Meeting notes and action items pushed into the CRM or the project tracker on their own
  • English and Spanish both handled — Houston work orders rarely arrive in only one

Chatbots & Customer-Facing Assistants

An assistant on your site or portal that answers from your documents, not the internet.

  • Grounded in your catalog, price sheets, policies, and manuals — with the source cited
  • Order status, tracking, and account questions answered from your live systems
  • Hands off to a person the moment it is out of depth, transcript attached
  • Guardrails on what it may discuss, quote, or promise on your behalf
  • Internal versions too: the help-desk assistant that has read every SOP you own

Vision & Photo Automation

The camera roll your crews fill up, turned into records.

  • Damage, delivery, and condition photos filed against the right load, job, or asset
  • Labels, VINs, serial plates, and pallet tags read from a phone photo instead of typed
  • Handwritten tickets and inspection forms digitized — the ones OCR alone gives up on
  • Photo evidence attached to the claim, the work order, or the inspection record automatically
  • Runs on-device or on your own server when the images shouldn't leave the building

Forecasting, Scoring & Anomaly Detection

The part of AI that isn't an LLM — ordinary models on your own history.

  • Demand and volume forecasts built from your order history, not an industry average
  • Late-shipment, no-show, and churn risk scored while there is still time to act on it
  • Duplicate invoices, billing anomalies, and rate discrepancies flagged before they're paid
  • Staffing and inventory plans that respect the seasonality you already live with
  • Judged against the boring baseline — if last month's number wins, we tell you that

AI Readiness & Data Foundation

What Lean Six Sigma is to process, this is to your data.

  • An honest inventory of what data you have and what state it is in
  • The gap between the AI use case you want and the data it needs
  • Quick wins ranked by payback, not by demo appeal
  • Data cleanup and plumbing before the model, not after the disappointment
  • A pilot scoped in weeks, with a measured baseline to judge it against
Advisory & aftercare

Not every AI job is a build

Some of the most useful weeks we spend on this produce a policy, a training session, or a number — not software. These are engagements in their own right.

AI policy & acceptable use

A short written policy your staff can actually follow: what may be pasted into a public chatbot, what may not, which tools are approved, and who to ask.

Shadow-AI audit

Which AI tools your team is already using, on what data, under whose terms. It is usually a longer list than management expects.

Training on the tools you already own

A working session for your team on the AI in the licences you already pay for, with prompts and checklists for their actual jobs rather than a generic webinar.

Tool & vendor selection

An honest read on the AI product a vendor is quoting: what it does, what it costs at your volume, and how much of it is a feature you already own.

Proof-of-concept bake-off

Two or three approaches run against your own documents and scored on the same sample, so the decision is made on measured accuracy instead of a demo.

Monitoring & support after launch

Accuracy tracked, spend capped and alerted, model versions pinned, and a migration plan for the day your provider retires the one you are on.

Any of these can be bought on its own. None of them obliges you to build anything with us afterwards — and if the honest finding is that you already own the tool you need, that is what the report will say.

Where it lands first

What this looks like in your industry

The same handful of patterns, pointed at the paperwork a particular business actually drowns in. These are the Houston operations where the payback shows up fastest.

3PL, freight & distribution

BOLs, PODs, rate confirmations, and status emails all day long. Extraction covers the partners who will never do EDI, and the exceptions still reach a person.

EDI & B2B logistics integration →

Manufacturing & wholesale

Purchase orders and acknowledgements keyed by hand, spec sheets nobody can search, and a quoting process that lives in one person's memory.

System integration →

Construction & field services

Photos, timesheets, and daily reports from the field turned into records — and voicemail from the job site turned into a work order before it is forgotten.

Mobile apps for field crews →

Fire & life safety

Inspection paperwork, panel programs, and deficiency lists read into records that produce an NFPA 72 report on demand instead of a folder of scanned PDFs.

Fire alarm programming & records →

Professional services

Contracts and engagement letters summarized, intake forms read, and follow-ups drafted from calls — with a person approving anything a client will see.

Customer intake & outreach →

Energy & industrial services

Field tickets, JSAs, and vendor invoices matched against the contract, and daily reports assembled from the documents the crews already send in.

Data entry & analysis →
How it runs

From one document type to something you rely on

Nothing here needs a platform contract or a year of discovery. The first useful result usually lands in weeks.

  1. 1

    Free assessment

    An hour on what your team retypes, what it costs you in time, and which of it a model can read reliably. You get a ranked list either way.

  2. 2

    Bake-off on your sample

    Two or three approaches scored against a real batch of your own documents, so the accuracy number you plan around came from your paperwork, not a vendor's.

  3. 3

    Pilot in production

    One queue, live, with validation rules and a review screen. Your team approves every call the model isn't sure about, and we watch the error rate together.

  4. 4

    Widen — or stop

    The pilot either beat the baseline or it didn't. If it did, the next document type is cheap. If it didn't, you spent weeks instead of a three-year commitment.

What you keep

Source code in your repository, prompts and validation rules in plain files you can read, API keys in accounts billed to you, and documentation written for whoever inherits it. If we deployed models on your own hardware, they keep running whether or not we are still involved — and if you would rather someone else maintain it, none of it is locked to us. That is the same deal as our custom software work.

What AI won't do

It won't fix a broken process — automating a mess just gives you a faster mess, which is why we sometimes recommend a process improvement pass before any model gets involved. It won't be right 100% of the time, which is why everything we ship has validation rules and a human veto on the calls that matter. And it won't replace your team: the goal is that people stop retyping and start reviewing.

It also isn't a strategy. "We need AI" is not a project; "stop hand-keying 400 invoices a month" is. We scope the second kind — small, measurable, and honest about the cases where the right answer is a boring script rather than a model.

Where it pairs well

  • Extraction for the trading partners who won't do EDI — the fax-and-PDF crowd
  • A Power BI dashboard tracking extraction accuracy and queue health
  • A legacy system that needs a modern front door before AI can reach it
  • The Control phase of a Lean Six Sigma project — automation that holds the gain
  • A mobile app that puts the camera and the review queue in the crew's hands
  • A cleanup pass on the data first, so the model isn't learning from duplicates
  • The integration layer that gets the extracted values into the ERP without a person in between
  • A feasibility study when the question is whether the thing can be built at all

Pilot first

One document type, one queue, a few weeks. A measured pilot beats a platform contract you're stuck with for three years.

Already on Microsoft?

AI Builder and Power Automate can carry the first workload inside the Power Platform licences you already pay for.

Needs an app around it

Review screens, queues, and audit trails are custom software — we build that side too, so the model isn't an orphan.

Fastest payback: logistics

BOLs, PODs, rate confirmations, and status emails all day long — 3PLs and distributors are where this earns its keep first.

Toolkit

What we work in

Model-agnostic on purpose. Which one runs a given workload is a configuration choice you can change later, not an architecture you're married to.

Claude API OpenAI / GPT API Azure OpenAI Ollama (self-hosted) Llama (open-weight) Qwen (open-weight) Python Retrieval-augmented generation Vector search & embeddings pgvector Structured extraction Function calling & tool use Model Context Protocol (MCP) OCR & document parsing Azure AI Document Intelligence Computer vision Speech-to-text & call transcription Twilio voice & SMS Forecasting & anomaly detection (scikit-learn) Power Platform AI Builder Power Automate + AI Evaluation harnesses & accuracy tracking Human-review workflows
Questions we get

Straight answers

Our data can't leave the building. Does that rule out AI?

No — it rules out the SaaS tools, which is a different thing. We deploy open-weight models like Llama and Qwen on your own hardware via Ollama: no per-token fees, no data crossing your network boundary, no vendor terms to renegotiate. We run local models in our own operation, so this is a setup we maintain daily, not a slide.

What does this cost, and where do we start?

Start with a pilot: one document type or one queue, fixed scope, a few weeks, and a measured baseline so you can judge it honestly. That is a fraction of what a platform contract costs, and it tells you whether the bigger build is worth doing before you commit to it.

How long before something actually works?

A document-extraction pilot typically shows real results in two to four weeks. The model is rarely the slow part — the plumbing into your systems, the validation rules, and the review workflow are where the engineering lives, and that is ordinary software work with ordinary timelines.

Do we need to hire an ML team?

No. Practical AI in 2026 is systems integration, not research — the models are built; the work is wiring them safely into your processes. What you keep is normal software: readable code, documentation, and monitoring your existing people or a modest support arrangement can run.

Our systems are old. Does that rule this out?

Old systems are our home turf. AI usually sits alongside the legacy core rather than inside it — reading the documents, files, and emails around it, and talking to it through the database, an API layer we add, or the same interfaces your staff use today. No rip-and-replace required.

What about hallucinations?

Real, and managed like any other failure mode. We constrain the model to extraction and drafting rather than open-ended answers, validate output against your business rules, route low-confidence cases to a person, and measure accuracy over time. Nothing we ship answers your customers unsupervised on day one.

Which model do you use — ChatGPT, Claude, or something else?

Whichever one wins on your workload. We build against an abstraction rather than a vendor, so a model is a configuration choice you can change later: Claude and GPT through their APIs or Azure OpenAI, and Llama- or Qwen-class open-weight models when the work has to stay on your hardware. Where it matters, we score two or three against the same sample of your own documents and let the numbers pick.

Can AI answer our phone or deal with voicemail?

Yes, and it is one of the fastest wins for a business that misses calls after hours. The realistic version takes the caller's details, answers what it can from your own information, and files a ticket or work order with everything captured — rather than pretending to be a person. Recordings and voicemail get transcribed, summarized, and attached to the right customer, in English or Spanish.

Do you build chatbots for our website?

We do, with the boring parts done properly: it answers from your catalog, policies, and manuals with the source cited, it can look up a real order status instead of guessing, it has explicit limits on what it may quote or promise, and it hands off to a person with the transcript attached the moment it is out of depth. A bot that invents a price is worse than no bot.

We already pay for Microsoft 365 Copilot. Do we need anything else?

Often not for the general assistant work — if you own the licences, use them, and we will happily train your team on them instead of selling you a second tool. What Copilot does not do is read your inbound documents into your ERP, answer from systems it cannot reach, or run a queue with validation and an audit trail. That gap is the part worth building.

How do you keep the running cost predictable?

We size the model to the job rather than defaulting to the largest one, cache and batch where the workload allows, and put a hard monthly cap with alerting on every API key we set up. You get the per-document cost measured during the pilot, so the monthly bill is something you approved in advance rather than discovered. Self-hosted models remove the per-token line entirely.

Is this going to replace our staff?

Not in anything we have built. The pattern that works is that the model does the reading and the typing while your people do the checking and the exceptions — the same headcount handling more volume, with the tedious part gone. If a vendor tells you a model will run a queue unsupervised, ask them who signs off on the mistakes.

Do you work with businesses outside Houston?

Yes. We are based in Houston and are glad to sit in your conference room anywhere in the Greater Houston area — Sugar Land, Katy, The Woodlands, Conroe, Pearland — and we also run these engagements remotely for clients across Texas and the rest of the US. Document automation in particular rarely needs anyone on site.

Get in touch

Send us the document your team retypes the most.

We'll tell you whether AI can read it reliably, what a pilot costs, and when the honest answer is a plain script.

Prefer to talk?

Call 832-598-8234 or email msco@stoneagesoftware.com. Houston, Texas — serving Houston, The Woodlands, Conroe, Sugar Land, Katy, Pearland, and the Greater Houston metro.

What happens next

We read your message, an engineer replies with questions or a straight answer, and if it's worth a call we book one. No drip campaign, no handoff to sales.

No obligation — we reply within one business day. This form sends your details to Stone Age Software LLC so we can respond to you. See our privacy policy for how we handle them. This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.