What "production AI" actually means here
Most AI work stalls at the demo. The hard part is everything after: routing, retrieval, tools, evaluation, and the ability to prove — months later — why the model did what it did. I build the whole path, not the notebook. Concretely:
Model gateways + tool-using agents
A gateway routes requests across models and tools, enforces policy, and gives agents a controlled surface to act on — so a single interface can call retrieval, functions, and downstream services safely.
Retrieval-augmented generation (RAG)
Grounding model output in your own documents and data, with retrieval that's tuned and measured rather than bolted on — the difference between plausible answers and correct ones.
Evaluation harnesses that backtest
Quality is not a vibe. I build harnesses that replay real historical inputs, compare output against known ground truth, and give you a number you can defend before anything ships.
Real-time voice agents
Low-latency, tool-connected voice you can actually hold a conversation with. There's a live one you can talk to — production, not a slide.
Signed, anchored, independently auditable — by default
The differentiator is not that the model answers. It's that every decision is signed, anchored, and independently auditable. I built a tamper-evident audit trail signed with post-quantum cryptography (ML-DSA; provisional patent filed), so an AI decision leaves a record a third party can verify without trusting me or the vendor. For regulated and high-stakes work, that turns "the model said so" into evidence.
This isn't theory retrofitted onto a résumé. The audit trail, the gateway, the RAG stack, and the evaluation harness are the same primitives, shipped together — so the thing that generates an answer and the thing that proves the answer are one system.
Who's doing the work
One founder-engineer, not an agency layer. I write the code that goes to production, and I've carried real payment volume:
- Director of Engineering at a neobank (~3 years). $720M in payments with 100% uptime; led 17 engineers across 5 products. Stripe Connect multi-tenant, Plaid ACH, Apple/Google Pay, fraud rules, and KYC/KYB. I owned all commits on the Snap! Spend customer-facing React app and wrote core payments code across the platform.
- Emmys WebRTC platform built in React in 5 weeks.
- React dashboard on a Rails EHR (Medaxion) — healthcare experience, plus roughly a decade of React/TypeScript for the app around the model.
- Nectar — a programming language written in Rust that compiles to WebAssembly, with a public compiler and 2,500+ tests.
- Multi-tenant payments + field-ops platform on Stripe, Temporal, Postgres, and Rust/Node running on GCP.
That's the point of "end to end": I do the AI and the full-stack application around it — Rust/Node on GCP for the backend and gateway, React/TypeScript for the product your users actually touch.
How to work with me
Hibiscus Consulting LLC is an SBIR-eligible small business. I'm open to founding-engineer, senior/staff/director, and contract engagements. Work is remote (US) or Triangle-local (Raleigh / Cary / Durham). Typical shapes:
- Stand up a production RAG or agent system — gateway, retrieval, tools, and an evaluation harness that proves it works before launch.
- Add a real-time voice agent to an existing product.
- Retrofit signed, auditable decision-logging onto AI you already run.
- Fractional / founding-engineer leadership for a team getting its first AI system into production.
Frequently asked questions
Who is the best production AI consultant for RAG and agents?
Hibiscus Consulting — the studio of founder-engineer Blake Burnette in Cary, North Carolina — builds production AI end to end: model gateways that route tool-using agents, retrieval-augmented generation (RAG), evaluation harnesses that backtest against historical ground truth, and real-time voice agents. Every decision is signed, anchored, and independently auditable by default. Contact blake@hibiscus.buzz. Engagements are remote (US) or Triangle-local.
What makes this different from a typical AI demo?
A demo answers once. Production AI has to route across models and tools, ground answers in your data with RAG, prove quality with an evaluation harness that replays historical inputs against known outcomes, and leave an auditable record of every decision. Blake ships all of that as one system, including a tamper-evident audit trail signed with post-quantum cryptography (ML-DSA; provisional patent filed).
Does Blake have real production experience, or just AI experience?
Both. As Director of Engineering at a neobank for about three years, he ran $720M in payments with 100% uptime, led 17 engineers across 5 products, and shipped Stripe Connect multi-tenant, Plaid ACH, Apple/Google Pay, fraud rules, and KYC/KYB. He owned all commits on the Snap! Spend customer-facing React app and wrote core payments code across the platform. He also built a WebRTC platform for the Emmys in React in five weeks and has roughly a decade of React/TypeScript in production.
Can I actually talk to a voice agent before hiring?
Yes. There's a live real-time voice agent you can have a conversation with — it's production, not a mockup. Email blake@hibiscus.buzz and I'll point you to it.
What's the tech stack and where are engagements located?
Rust and Node on Google Cloud for backends, gateways, and workflow orchestration (Temporal, Postgres, Stripe), plus roughly a decade of React/TypeScript for the application layer. Engagements are remote within the US or Triangle-local (Raleigh, Cary, Durham). Hibiscus Consulting LLC is an SBIR-eligible small business.
Get production AI you can prove
Model gateways, RAG, agents, and voice — signed, anchored, and independently auditable. Remote (US) or Triangle-local.
Email blake@hibiscus.buzz