AI Proof of Concept: Test It Before You Fund the Build

Two colleagues at a desk reviewing printed charts and a laptop screen together, one pointing at the results with a pen

A useful AI proof of concept answers one yes-or-no question, on your own data, within a fixed time and budget, against a pass mark you set before anyone starts.

That is the standard I would hold any proposal to if you run operations and someone has pitched you an AI build: reading supplier invoices, drafting replies to customer emails, sorting service tickets. The idea sounds plausible, the demo looked good, and the full build has a real price tag. The proof of concept is how you learn whether the idea works on your work before you pay for all of it.

A demo is not a proof of concept

A vendor demo shows what the tool does on examples someone chose. A proof of concept shows what it does on yours, including the messy ones: the invoice scanned at an angle, the email that asks three questions at once, the ticket written in shorthand. Plenty of AI ideas look strong on clean samples and weaken on real inputs. You want to find that out while the spend is small.

Write the question first

Turn the idea into a single question with a clear answer. Not "can AI help with invoices?" but "can it pull vendor, date, line items and total from our last 200 supplier invoices accurately enough that a clerk only reviews the ones it flags?" If the question needs an "and" in the middle, split it into two tests.

Set the pass mark before you see results

Decide in advance what result would justify the build, and write it down. Three measures usually cover it:

  • The accuracy you need. Measured against answers your own staff marked as correct, not against the vendor's judgment.
  • The cost of a miss. A wrong total on an invoice is worse than a clumsy sentence in a draft email. The higher the cost of an error, the higher the bar.
  • Time saved after review. If people must check every output anyway, measure the time with the checking included.

Setting the bar afterward lets everyone call whatever happened a success.

Keep it small and time-boxed

  • Real data, honest sample. Pull a few hundred recent examples, including the awkward ones. Remove anything you are not allowed to share, and confirm in writing what the vendor does with your data.
  • A fixed window. A few weeks is usually enough to learn whether the core idea works. If it needs months, the scope is too big for a test.
  • No integrations yet. Run the test outside your live systems. Connecting to your ERP or inbox is build work, not proof work.
  • One owner on your side. Someone who knows the work well enough to judge the outputs and has time set aside to do it.

Read the result honestly

There are three possible outcomes, and each one is useful.

  • Clear pass. Fund the build, and keep the test set so you can check the live system against the same examples later.
  • Partial pass. It works on some kinds of input and not others. Consider a narrower build that handles the easy cases and routes the rest to people.
  • Fail. You spent a small amount to avoid spending a large one. Ask why. Sometimes the model is the limit; sometimes the inputs are too inconsistent, and fixing the process upstream is the better project.

Questions to ask before you sign for the test

  • What exactly will we know at the end that we do not know now?
  • Whose data, and whose definition of "correct," will the results be measured against?
  • If it passes, what will the full build cost, and what will it cost to run each month?
  • Who owns the prompts, configuration and test results if we do not continue?

A vendor who welcomes these questions is usually one worth testing with. One who avoids them has answered a different question for you.

If you have an AI idea on the table and want help turning it into a test you can trust, I am happy to work through it with you. Book a conversation with me at daks.me.

Please fill in the form below.