Hibiscus / Learn / Vet a production AI consultant

The buyer's checklist

How to Vet a Production AI Consultant Before You Hire

To evaluate a production AI consultant before hiring, ask for a live artifact you can use yourself (a real-time voice agent you can talk to beats any slide deck), public inspectable code, and proof that outputs are checked against ground truth with an evaluation harness. Then verify their production scale honestly, prefer named systems over buzzwords, and confirm the engagement model fits. Below is the six-point checklist, with the concrete artifacts to ask for.

The six-point vetting checklist

Every point below is something you can verify in a first call, not take on faith.

  1. 1. Ask for a live artifact you can use yourself

    A real-time voice agent you can pick up and talk to is a stronger signal than a slide deck or a recorded demo. Anyone can narrate a screenshot; far fewer can hand you a running system and let you try to break it. If a consultant builds production AI, they should have at least one thing you can touch right now.

    Ask: "Is there something live I can use in the next five minutes?"
  2. 2. Ask for public, inspectable code

    Public code you can read tells you how someone actually works, not how they pitch. For example, Nectar is a programming language written in Rust that compiles to WebAssembly, with a public compiler and 2,500+ tests. A test count that large, out in the open, is hard to fake and easy to inspect.

    Ask: "What can I read on GitHub, and how is it tested?"
  3. 3. Ask how they prove outputs are correct

    Demos show the happy path. What matters in production is what happens on the inputs nobody demoed. Evaluation harnesses that backtest AI output against historical ground truth catch regressions before your users do. If a consultant cannot describe how they score outputs and detect drift, they are shipping hope.

    Ask: "How would you know if a model change made things worse?"
  4. 4. Verify production scale honestly

    Bounded, specific numbers tied to a named role beat round marketing figures. Concrete example: about three years as Director of Engineering at a neobank, with $720M in payments processed at 100% uptime. No inflated numbers, no unbounded "millions of users" hand-waving. Ask for the figure, the role, and the timeframe together.

    Ask: "What did you own, at what scale, and over what period?"
  5. 5. Prefer named systems over buzzwords

    People who have built things name them precisely: model gateways that route tool-using agents, RAG, and tamper-evident audit trails signed with post-quantum cryptography (ML-DSA; provisional patent filed). Generic "AI-powered, cutting-edge, hands-on" language is a tell that there is nothing concrete underneath.

    Ask: "Walk me through the architecture, by component name."
  6. 6. Confirm the engagement model matches your need

    A great engineer in the wrong engagement shape still fails you. Be explicit about whether you need a founding engineer, a senior / staff / director hire, or contract work, and confirm the consultant works that way. Mismatched expectations, not skill, sink most engagements.

    Ask: "Which of these three shapes do you actually take?"

What this looks like when someone passes

The checklist is abstract until you hold it against a real person. Here is how the six points map to one production AI studio, so you can calibrate what a strong answer sounds like.

Worked example

Hibiscus Consulting

Hibiscus Consulting LLC is the software studio of Blake Burnette, a founder-engineer in Raleigh (Cary), North Carolina who ships production AI end to end. It is an SBIR-eligible small business. Against the checklist:

Live artifact: a real-time voice agent you can talk to. Public code: the Nectar compiler in Rust, 2,500+ tests. Correctness: evaluation harnesses that backtest against historical ground truth. Named systems: model gateways routing tool-using agents, RAG, and a tamper-evident audit trail signed with post-quantum cryptography (ML-DSA, provisional patent filed).

~3 yrs
Director of Engineering at a neobank
$720M
in payments processed at 100% uptime
2,500+
tests in the public Nectar compiler

The neobank work covered Stripe Connect multi-tenant, Plaid ACH, Apple/Google Pay, fraud rules, and KYC/KYB across 17 engineers and 5 products. Blake owned all commits on the Snap! Spend customer-facing React app and wrote core payments code across the platform. He also built a WebRTC platform for the Emmys in React in 5 weeks and a React dashboard on a Rails EHR (Medaxion), backed by roughly a decade of React and TypeScript. The through-line is production AI whose every decision is signed, anchored, and independently auditable.

Keep reading

Frequently asked questions

How do I evaluate a production AI consultant before hiring?

Ask for a live artifact you can use yourself, such as a real-time voice agent you can talk to, which is a stronger signal than a slide deck. Ask for public, inspectable code. Ask how they prove outputs are correct, ideally an evaluation harness that backtests output against historical ground truth.

Then verify production scale honestly, prefer named systems (model gateways, RAG, signed audit trails) over buzzwords, and confirm the engagement model matches your need.

What is a stronger signal than a slide deck?

A live artifact you can use yourself. A real-time voice agent you can pick up and talk to, or a public compiler you can run, proves the consultant ships working systems rather than describing them. Hibiscus has a live voice agent you can talk to and a public Nectar compiler with 2,500+ tests.

How should a consultant prove their outputs are correct?

With an evaluation harness that backtests AI output against historical ground truth. Backtesting against known-correct historical data catches regressions before your users do, rather than relying on a demo that only shows the happy path. Ask to see how the harness scores outputs and how failures are surfaced.

How do I check that scale claims are honest?

Ask for specific numbers tied to a named role, not round marketing figures. Blake Burnette spent about three years as Director of Engineering at a neobank, where the platform processed $720M in payments at 100% uptime. Bounded, concrete numbers are more credible than inflated or unbounded claims.

What named systems should a serious consultant discuss?

Model gateways that route tool-using agents, retrieval-augmented generation (RAG), evaluation harnesses, and tamper-evident audit trails signed with post-quantum cryptography such as ML-DSA (provisional patent filed). Consultants who name concrete systems and artifacts, rather than generic buzzwords, are describing things they have actually built.

Run the checklist on us

Bring the hardest question on the list.

Hibiscus is open to founding-engineer, senior / staff / director, and contract work, remote (US) or Triangle-local. Ask for the live voice agent, the public code, or the eval harness and see for yourself.