Business Automation

10 Questions to Ask Before Hiring an AI Consultant

Ten questions that separate a real AI consultancy from a rebranded chatbot shop — the answers that should worry you, and how we answer them ourselves.

By Isaac, Founder, Visione Edge12 min read
Low-angle chessboard on dark stone: two identical knights under a warm ivory spotlight — one solid carved wood casting a real shadow, one hollow glass casting none.

What should you ask an AI consultant before you sign?

Ask ten questions before hiring an AI consultant: live deployment, agent or chatbot, acceptance criteria, who does the work, post-deployment support, audit logs, what they won't build, data handling, past failures, and pricing structure. Good answers are specific and verifiable. Worrying answers are fluent, confident, and impossible to check. The checklist pairs each question with a good answer and a worrying one, so you can hear the difference in the room.

The stakes are not hypothetical. Gartner predicts that "over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls." Most of those cancellations start the same way: a buyer signed without asking questions the vendor could not answer well.

One disclosure before the list. We are Visione Edge, an AI consulting studio — one of the firms you might be evaluating. So read this as both a checklist and a commitment: at the end of the article, we answer the hardest of these questions about ourselves, on the record.

#The questionA good answer sounds likeAn answer that should worry you
1Can you show me a live deployment, not a demo?"Yes — here is one running, and here is what we can't show because it's under NDA. I'll tell you plainly which is which."A polished deck, a scripted video, or "we're just ramping up delivery."
2Is this an agent, or a chatbot with a new label?A plain explanation of what the system decides on its own — and which parts stay scripted."Agentic" used in every sentence, but no example of a single decision the system makes.
3How will we know it works before final payment?Written acceptance criteria: tests on your data, pass rates, and who runs them."You'll see it in the results," or a demo run only on the vendor's own examples.
4Who actually does the work?Named people and roles, with the split between senior time, junior time, and subcontractors.The person selling can't say who delivers; the team "will be assigned later."
5What happens after deployment?A support window, monitoring, an escalation path, and a handoff plan — written into the statement of work (SOW)."It runs itself," or support exists only as a second contract at new prices.
6Who holds the audit log?"You do. We set up logging you can read without us in the room."Logs live in the vendor's systems and leave when the vendor does.
7What won't you build?Specific projects they declined and why — usually thin data, no owner, or no measurable outcome."We can automate anything." Nothing is ever out of scope.
8Where does our data go?A data-flow answer: which model providers, what gets retained, what the data agreement says."It's all secure," with no named subprocessors (the third parties that touch your data) and no retention answer.
9What does failure look like? Have you had one?A real story, with what changed in their process afterwards."Our projects don't fail."
10How do you price, and what changes the price?Fixed scope with stated assumptions, or capped day rates — with discovery priced separately.One confident number before any discovery, or an open-ended retainer with no exit.

Question set and answer patterns: Visione Edge analysis, July 2026. Formats like this exist across the vendor-selection space; the test is whether the firm in front of you survives its own row.

Why is "show me a live deployment" the first question?

Because a demo is a promise and a deployment is evidence. A demo shows the happy path the vendor rehearsed. A live deployment shows what survived real users, real data, and real edge cases. A firm that builds production systems can always show you something real — and when client work is confidential, it can still tell you plainly what is a demo, what is anonymized, and what is under NDA.

Accept NDA'd reference calls, anonymized architecture walkthroughs, and public demos labeled as demos. What you should not accept is the substitution trick: a rehearsed demo presented as if it were production evidence.

A firm that will not show you a live deployment is selling you a deck, not a system.

How do you spot agent washing when hiring a consultant?

Ask what the system decides on its own, which parts are fixed, pre-written scripts, and to see one decision trace from a real log. A real agent shop answers in minutes. A rebranded-chatbot shop answers with vocabulary. This matters because renaming is now an industry practice with a name — and an analyst estimate attached.

From Gartner's June 25, 2025 press release, verbatim:

Many vendors are contributing to the hype by engaging in "agent washing" – the rebranding of existing products, such as AI assistants, robotic process automation (RPA) and chatbots, without substantial agentic capabilities. Gartner estimates only about 130 of the thousands of agentic AI vendors are real.

About 130, out of thousands. You do not need to audit a vendor's codebase to protect yourself from those odds. You need question 2 from the table, asked twice: once in the sales call, once to the engineer who would deliver. If the two answers differ, you have learned something the proposal would never have told you. Renamed technology is also one of the quieter reasons AI agent projects fail: the tool was never capable of the autonomy the contract assumed.

To be clear, a scripted workflow or a plain chatbot is often the right, cheaper answer. The problem is not the technology. The problem is paying agent prices for it without knowing.

What are acceptance criteria, and why do they belong in the contract?

Acceptance criteria are pre-agreed tests the system must pass — on your data, at a stated pass rate, run by a named party — before you make final payment. They convert "it works" from the vendor's opinion into a measurable event. If you have ever held back a builder's final payment until inspection, you already know how this works. AI procurement just forgets to do it.

Here is why asking is such an efficient filter. LangChain's State of Agent Engineering survey (1,340 respondents, fielded November 18 to December 2, 2025) found that 52.4% of organizations report running offline evaluations on test sets — while 89% have implemented some form of observability. Translation: nearly everyone can watch their agents, but only about half systematically test them. A vendor with real acceptance tests is, on that data, already in the more disciplined half of teams building agents.

Three things to demand in writing:

If the vendor cannot produce tests, a threshold, and a named runner, question 3 has answered itself. Demand all three in the statement of work — not in a follow-up email after signing.

Freelancer, agency, boutique, or in-house: which should you choose?

Match the structure to the problem's size and lifespan. A freelancer fits one well-defined workflow. An agency fits volume and many integrations. A boutique studio fits systems that need architecture and production standards on a fixed budget. An in-house hire fits companies where AI is core strategy for years. The rows most selection guides skip — accountability after deployment, and who keeps the evidence — are where these options genuinely differ.

CriterionFreelancerAgencyBoutique studioIn-house hire
Engagement shapeHourly or small projectProject or retainerScoped build plus support windowSalary
Strongest whenOne well-defined workflowMany integrations, volume deliveryProduction system, fixed budget, high standardsAI is core to the business long-term
Who does the workThe person you hiredMixed seniority; sometimes subcontractedSmall senior teamYour employee
Post-deployment accountabilityUsually ends at handoffPer contract — often a separate retainerShould be written into the SOW: support window and escalation pathPermanent — but you carry all of it
Who holds the audit logOften nobody; ad hocTypically the agency's stackShould be handed to you at go-liveYou, by default
Watch out forOne person is the whole teamThe pitch team is not the delivery teamSmall bench — confirm availabilityMonths, not weeks, to hire; scarce talent

Visione Edge analysis, July 2026. Disclosure: we sit in the boutique column, and we wrote the criteria we believe we win on. Test us on them — and test everyone else too.

What does it look like when a consultancy skips its own checks?

It looks like the Deloitte Australia case: a AU$440,000 government report shipped with AI-fabricated citations, corrected only after an outside academic caught them, followed by a partial refund. It is the cleanest documented example of what this article's questions exist to prevent — a vendor's internal verification gate failing on an AI-produced deliverable, and the client running quality control after paying.

The facts, from the public record:

FactDetailSource
ContractAU$440,000 (US$291,245), Australia's Department of Employment and Workplace Relations (DEWR)The Register, Oct 6, 2025
DeliverableIndependent review of the Targeted Compliance Framework, commissioned December 2024, published July 2025The Register
What shippedCitations to non-existent academic works, phantom footnotes, and a made-up quote attributed to a Federal Court judgmentFortune, Oct 7, 2025; The Register
Who caught itDr. Christopher Rudge, University of Sydney — after publicationFortune; AI Incident Database #1193
The AI disclosureCorrected version discloses a "generative AI large language model (Azure OpenAI GPT-4o) based tool chain" was used to fill "traceability and documentation gaps" — disclosed only after the errors surfacedThe Register
The refundMore than AU$97,000 (about US$63,000) — the final installment — confirmed by a DEWR spokespersonCFO Dive, Oct 21, 2025 (amount, DEWR confirmation); The Register (final installment)

Note what the failure was not. It was not "AI wrote the report." Tools are not the scandal. The failure was process: an AI-produced draft went out under a Big Four letterhead without a human verification pass rigorous enough to catch invented sources — and without telling the client a model was involved at all.

What follows is our analytical counterfactual — our reading of the public record, not a finding from any investigation. Three questions from the table above would likely have surfaced the problem before publication, or deterred it entirely:

  1. "Who actually does the work?" (question 4) — asked precisely, this becomes: does a human expert review every citation before delivery? A client who asked it up front would have forced the tool-chain disclosure that eventually appeared, months late, in the corrected version's methodology section.
  2. "How will we know it works before final payment?" (question 3) — an acceptance gate for a research report is citation verification. DEWR effectively ran acceptance testing after publication, through a University of Sydney academic it never hired.
  3. "Which AI tools touch our deliverable, and where is that disclosed?" (question 8's process half) — the corrected report answered this question honestly. The original did not, and the gap between those two documents is what the refund paid for.

The buyer here was a national government with a procurement department. If it can skip these questions, so can you — which is the argument for writing them down.

How should an AI consultant price the engagement?

Expect the pricing structure to be legible before any number is: what is fixed, what varies, which assumptions the quote depends on, and what discovery costs separately. A vendor who cannot explain what changes the price is pricing your uncertainty, not your project. And rates vary too much by scope for any one anchor to be honest — so treat every published band as a rough reference, never a quote.

As one such reference — a single published guide's bands, not a benchmark — AI Essentials (updated June 24, 2026; retrieved July 4, 2026) puts independent AI consultants at $150–$350 per hour, project work at $5,000–$25,000, and monthly retainers at $2,000–$8,000.

The worrying answers repeat a pattern: a confident total before discovery, or an open-ended retainer with no defined end state. Both signal a quote built to absorb unknowns that you, not the vendor, will end up paying for.

The full cost question — build cost versus run cost, and what actually blows AI budgets — deserves its own math, and we've done it in our breakdown of what AI automation really costs a small business.

When should you not hire an AI consultant at all?

Skip the consultant — for now — if the process you want to automate is not documented, if no one on your team will own the system after handoff, if an off-the-shelf subscription already covers the workflow, or if the budget only covers the build and nothing after it. A good firm will tell you this in the first call. A worrying one will take the engagement anyway.

There is no shame in the "not yet" answer. Automating an undocumented process just automates the confusion, and a system nobody owns internally degrades from the day the consultant leaves. If you are still one step earlier — deciding whether automation pays back at all in a business your size — start with whether AI automation is worth it for your business and come back to this checklist when the answer is yes.

How do we answer our own questions?

We answer the three hardest questions from our own table in writing — live deployment, refused work, and post-deployment ownership — so you can compare our answers against the "worrying" column before you ever talk to us. The rest we would rather answer live, against your actual project, where vague answers have nowhere to hide.

"Can you show me a live deployment, not a demo?" Our production client work is under NDA — we can discuss it in specifics on a call, not on a webpage. What we can show you publicly, today, is the Visione Flow booking agent demo running live on our site. And here is the distinction, stated plainly: that demo is a scripted demonstration of the architecture we build with — it is not a client system, and we will never present it as one. We hold ourselves to the same rule we gave you above: you will always know which is which.

"What won't you build?" Anything without a measurable outcome, an internal owner, or the data to support it. We negotiate scope, timelines, and budgets — never architecture, depth, or standards. If your project only fits the budget by cutting testing or logging, we will decline it and say why.

"What happens after deployment?" A defined support window, monitoring you can see, an escalation path with names on it, and audit logs that are yours — in your systems, readable without us. That is written into the SOW, not promised in the sales call.

Run these questions on us — 30 minutes, no pitch. Bring the table. Ask them in any order. If any answer worries you, that is the system working.

Sources

  1. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 — Gartner, 2025-06-25
  2. State of Agent Engineering — LangChain, 2025-12
  3. Deloitte refunds Aussie gov after AI fabrications slip into $440K welfare report — The Register, 2025-10-06
  4. Deloitte was caught using AI in $290,000 report to help the Australian government crack down on welfare after a researcher flagged hallucinations — Fortune, 2025-10-07
  5. Incident 1193: Purportedly Taxpayer-Funded Deloitte Report for Australian Government Contains Alleged AI-Generated Citations and Fabricated Legal Quote — AI Incident Database, 2025-08-22
  6. Deloitte refunds over $60K for report with AI errors, Australian government says — CFO Dive, 2025-10-21
  7. How Much Does It Cost to Hire an AI Consultant for My Small Business? — AI Essentials, 2026-06-24

Book an architecture call — 30 minutes, no pitch