← Back to blog

AI Snake Oil or the Real Thing?

The questions behind the answers

A client once told me they had built an agent that closed their books at the end of every month.

When we walked through it together, the agent became something smaller. It was a promising set of prompts and scripts that worked when the inputs arrived in the expected form.

I thought it was a strong experiment that needed another step before anyone relied on it. They thought it was ready.

A few months later, they called. The agent had been producing incorrect results, and nobody could say when it started drifting.

I have sat through more software demos than I can count. Increasingly, the AI product gets smaller as the questions get better.

That does not always make it a bad product. It may still be worth buying. But it is not the product we met at the beginning of the conversation.

Most AI snake oil starts with something real. The claims just get ahead of it.

The purpose of this guide is not to prove that an AI product is good or bad. It is to help you make the claim specific enough to evaluate. The ten questions below do that, whether you are choosing a tool for yourself or considering a company-wide investment. A $20-a-month personal tool may require a 20-minute test. A system with access to customer data or permission to act may require a pilot, legal review, and a named owner. The questions are the same. The depth of the answers should match the consequences.

The questions are becoming familiar. That is good. It also means the answers are getting better rehearsed. People know to say "human in the loop," "enterprise security," "continuous improvement," and "measurable ROI." Those answers sound responsible. They are also where the real questions begin. Which human is in the loop? When do they intervene? What can they see? Who is accountable when they miss something? How much time does all of this take?

What exactly is being promised?

Words such as "agent," "copilot," and "automation" are now used so broadly that they tell us almost nothing by themselves.

Ask what the product actually does. What task does it perform? Who uses it? What outcome does it produce? Where does the work begin and end? At what point does a person take over?

If the product is called an agent, ask what decisions it can make and what actions it can take without approval.

A credible answer usually gets narrower as it becomes more specific. That is a good sign.

Look for: One task, one user, a defined outcome, and a clear boundary.

Be cautious when: The language keeps changing, or the product seems to become whatever the conversation requires.

Can it prove itself on your work?

A polished demonstration tells you that the vendor can produce a polished demonstration. It does not tell you how the product will perform with your data, your people, your processes, or the exceptions that make your work difficult.

Ask to see the product operate on a representative sample of the real work. Include ordinary cases, difficult cases, and a few situations where the correct answer is unclear.

For an individual, this might mean testing the tool against documents or tasks you know well. For a company, it means a limited evaluation using real workflows, real constraints, and clearly defined success criteria.

Do not make the test unnecessarily hostile. You are not trying to trap the product. You are trying to learn what it needs to work.

Look for: A willingness to test the product with representative inputs, users, and constraints.

Be cautious when: The conversation keeps returning to benchmarks, prepared demos, or customer testimonials.

What would disprove the claim?

Before the test begins, ask what result would count as failure. This is harder than it sounds.

Once people become invested in an AI project, they can explain away nearly every disappointing result. The prompt needs work. The data was unusual. Users need training. The model will improve. The workflow needs to be redesigned.

Any one of those explanations may be true. Together, they can make a claim impossible to disprove.

Define the minimum acceptable result before seeing the outcome. Decide what will be measured, how it will be measured, and what would cause you to stop or reconsider.

A useful evaluation can end with "not yet," "not for this workflow," or "not at this cost."

Look for: Success and failure criteria established before the evaluation.

Be cautious when: Every miss becomes another reason to continue testing.

What happens when it is wrong?

An accuracy percentage is not a failure plan.

A system can be 95 percent accurate and still be dangerous if the remaining 5 percent includes consequential, expensive, or difficult-to-detect errors. A typo in an internal summary and an incorrect decision about a customer may have the same accuracy score. They do not have the same risk.

Ask what failure looks like in practice. Can users see when the system is uncertain? Are mistakes easy to detect? Can an action be reversed? Who is affected? What happens after an error is discovered?

The important measure is not only how often the system is wrong. It is the combination of frequency, visibility, and consequence.

Look for: A concrete failure story, clear escalation rules, and a recovery path.

Be cautious when: The answer begins and ends with an accuracy number.

How much human work is hidden?

Many AI products do not eliminate work. They move it.

Someone prepares the data. Someone writes or adjusts the prompts. Someone checks the result. Someone handles exceptions. Someone explains the mistake. Someone keeps the process working when the model changes.

That work may be worthwhile. It still needs to be counted.

"Human in the loop" is not a complete answer. Ask which human, doing what, how often, and with what level of expertise.

Pay particular attention to review work. Reading an AI-generated result carefully enough to catch a subtle error can take as long as producing the result in the first place.

I see this most often with internal IT teams. One company ran a hackathon that produced dozens of good ideas, several of which addressed expensive operating problems. Every one of them needed a request reviewed, code checked, and an integration approved by the same small IT group that already had a full queue. Within weeks, the IT team was swamped. The initiatives eventually stalled, not because the ideas were bad, but because nobody had counted the review work.

Look for: Setup, review, correction, and escalation work included in the operating model.

Be cautious when: The product is described as autonomous until the conversation turns to failures.

Does the value survive real use?

AI return on investment calculations often begin with time saved. Ten minutes saved per task, multiplied by the number of employees and their hourly cost, produces a wonderfully confident number.

But saved time is not the same as captured value.

Will people use the tool? Is the output good enough to use? How much time is spent reviewing and correcting it? Does the saved time become additional capacity, better work, or lower cost? Or does it simply disappear into an already crowded day?

For individuals, the value might be speed, access to a new capability, or reduced frustration. For companies, the benefit needs a path into the operating model. If nothing changes after the time is saved, the financial return may remain theoretical.

We have seen the other outcome, too. One client cut the cost of a back-office pipeline by 85 percent and reduced a process that took weeks to one that took days. We agreed what would be measured and how before anything was built. The project was designed around moving those numbers.

Look for: Adoption, output quality, rework, and the use of freed capacity included in the calculation.

Be cautious when: ROI is presented as hours multiplied by hourly cost.

How does the human work change?

Counting the work is one thing. The other question is what the job feels like afterward.

The usual question is whether AI will replace a person. The more immediate question is how it changes that person's job. Does the user gain time or a useful new capability? Do they make better decisions? Or do they inherit more monitoring, exception handling, and accountability?

A system can reduce the effort required to perform a task while making the job around that task more fragmented and difficult.

This matters especially when AI separates action from understanding. A person may remain accountable for an outcome while having less visibility into how the result was produced.

That is not automatically a reason to reject the tool. It is a reason to design the work, not just install the software.

Look for: A clear description of the user's new responsibilities and how the workflow changes.

Be cautious when: The human contribution is treated as either unnecessary or free.

What keeps it working?

The demo is the beginning of the cost, not the end.

AI systems depend on changing models, data, integrations, policies, and workflows. A process that works today may drift as any of those pieces change.

Ask who will own the system after launch. Who monitors quality? Who updates it when the underlying model changes? Who handles new failure patterns? Who supports users? What expertise is required, and what will it cost?

If the answer is your team, make sure your team knows that.

A successful pilot can still become an expensive production system if the decision never included its ongoing needs.

Early on, we helped a company build an AI workflow for internal operations. We tested it, strengthened it, and left the internal team able to keep improving it. Later, leadership eliminated the team. The system kept running, but no one remained to spot new failure patterns or adapt it as the work changed.

Look for: Named ownership, an operating process, and a realistic maintenance cost.

Be cautious when: The team that made the demo successful disappears after launch.

Where do the data and decisions go?

A link to a security page is not the same as an understandable answer.

Ask what data the system receives, where it is processed, how long it is retained, and whether it is used to train or improve any model. Ask who can access the data and which permissions the product inherits. If the product takes action, ask what authority it has and how it records those actions.

For a company, these questions should include vendors and subprocessors, not just the product you are buying. For an individual, the same principle applies. If you would not publish the information, understand what happens before you paste it into a tool.

A good answer should be clear enough that someone in the room can repeat it accurately afterward.

Look for: Plain language about sources, permissions, retention, control, and authority.

Be cautious when: Simple questions produce a stack of links but no direct answer.

What are its limits, and how do you leave?

Ask what the product should not be used for. A thoughtful vendor, adviser, or internal champion should be able to answer this comfortably. Every useful technology has boundaries.

Then ask what happens if you decide to stop. Can you export your data in a usable form? Can you recover prompts, configurations, histories, or workflow definitions? What happens to your data after termination? How much help will they provide during the transition?

If leaving means rebuilding your data, history, and workflows from scratch, that cost belongs in the original decision.

This is especially important with AI products because much of the useful knowledge may accumulate inside configurations, evaluations, and interaction histories that are difficult to move.

Look for: A real list of non-uses, clear stopping conditions, and a practical offboarding path.

Be cautious when: The product does everything, failures are always temporary, and leaving has never come up.

The answer is only the beginning

No single weak answer proves that a product is snake oil.

Some questions will not have perfect answers. New products have gaps. Small vendors have limits. Even strong teams are still learning how to operate these systems.

What matters is the pattern. Do the claims become clearer as you ask questions, or do they keep moving? Does uncertainty become visible, or is it covered with confident language? Does the full cost appear, or does the difficult work quietly move to you?

Good products usually become easier to describe as the questions get harder. The claim may get smaller, but it also becomes more believable. That is the point of these questions. Not to eliminate uncertainty, but to find out whether the people across the table understand it.

A quick version for your next conversation

Before adopting an AI product, ask:

  • What exactly is being promised?
  • Can it prove itself on our work?
  • What would disprove the claim?
  • What happens when it is wrong?
  • How much human work is hidden?
  • Does the value survive real use?
  • How does the human work change?
  • What keeps it working?
  • Where do the data and decisions go?
  • What are its limits, and how do we leave?

Do not listen only for the answer. Listen for whether the answer holds up to the next question.

Share this

Ron Davis

Founder

Three decades building enterprise platforms. Started Joust to close the gap between strategy decks and the work they're supposed to change.

LinkedIn

← Back to all posts

Get in touch