Skip to content
02 · Assistants that know your business, not the internet

AI agents & copilots

A general-purpose chatbot is a demo. An agent that quotes your pricing policy correctly, cites the contract clause it came from, and escalates when it isn't sure — that's a system. We build the second kind.

Answers grounded in your own documents, with citationsMeasured accuracy against a real test set, not vibesGuardrails and refusal behavior you definePer-conversation cost you can forecast
A reviewer comparing a printed document against a laptop screen

What you get

Retrieval-grounded agents and internal copilots built on your documents, your data, and your rules — with evaluation harnesses so you can prove they work.

01A retrieval layer that actually retrieves

Chunking, embedding, and reranking tuned to your document shapes — contracts read differently than ticket threads. We test retrieval quality separately from generation quality.

02An evaluation harness

A graded test set built from real questions your team asks. Every prompt or model change runs against it, so you know whether a 'small tweak' made things worse.

03Tool use, scoped tightly

Agents that can look up an order, draft a reply, or open a ticket — with permissions, rate limits, and a dry-run mode before anything writes to a system of record.

04Escalation the ops team trusts

Confidence scoring, refusal paths, and handoff to a person with full context attached. The fastest way to kill an assistant is to let it guess in public.

05Model portability

We abstract the provider. When a cheaper or better model ships — and one will — you switch with a config change and a re-run of the eval suite.

Engagements

Shapes this usually takes

Most engagements take one of these three shapes. Which one fits depends on how well-defined the problem already is — tell us the situation and we'll say which, and what it would take.

3 weeks

Agent prototype

One agent, one corpus, one clear job. Includes the eval set so you can judge it honestly.

Contact us
8–14 weeks

Production agent build

Full retrieval pipeline, tool integrations, guardrails, observability, and rollout to real users.

Contact us
Monthly

Agent operations

Eval maintenance, prompt and model tuning, corpus refresh, and cost optimization.

Contact us
Tools & stack

What we reach for

Boring where boring works, current where it matters. We pick for what your team can maintain after we leave — not for what looks good in a case study.

Two developers sharing one screen while working through code together
Claude, GPT, Llama, Mistral
pgvector, Pinecone, Weaviate
LangGraph, custom orchestration
Ragas, promptfoo, custom evals
OpenTelemetry, Langfuse
SSO, RBAC, audit logging

Common questions

Will our data be used to train someone's model?

Not on our watch. We deploy against enterprise endpoints with zero-retention terms, or self-hosted open-weight models when the data can't leave your perimeter. We'll put it in the contract.

How accurate is accurate enough?

It depends on what happens when it's wrong. We set the bar with you before the build — and if the honest answer is that the task needs 99.9% and retrieval can only get you to 92%, we say so rather than shipping it.

Can it work on documents that aren't digitized well?

Usually. Scanned PDFs, faxes, and photographed forms go through an OCR and structuring pass first. Quality varies, and we'll benchmark yours during scoping.

Talk it through

Is ai agents & copilots the right move for you?

Send us the situation. We’ll come back with a read on whether this is the right service, a rough range, and what we’d want to learn first.

We use your details to reply to this request only. No sequences, no list.

Taking two new engagements this quarter.