AI for Software Discovery: Where It Helps and Where It Fails

ยท 7 min read

AI is good at the slow, mechanical part of discovery: reading a pile of decks, transcripts and notes, and turning them into a first draft of goals, users, journeys and requirements. It is bad at knowing what nobody wrote down, and it can state invented details with complete confidence. Used well, it gives you a draft in minutes that a person then checks line by line against the sources; the decisions stay with people.

Where AI helps in discovery

Most discovery time goes into finding, sorting and restating information that already exists somewhere. That is where a language model earns its place.

Here is the kind of structure that a draft should end up in: users with their goals and pain points, and journeys that name those users. The project is invented (it is called Fieldwise, from the public example report):

The User experience section of the example report for the invented Fieldwise project: three actors with goals and pain points, two user journeys and two user flows

Where it fails: invention, missed context, over-confident text

The failures are predictable, which makes them manageable.

Invention. NIST's generative AI profile calls this confabulation: "the production of confidently stated but erroneous or false content" (NIST AI 600-1). In discovery it looks like a metric with a target nobody set, a stakeholder's job title filled in from a guess, or an acceptance criterion with a precise threshold that appears in no source.

Missed context. The model only knows what you give it. The political reason a previous project failed, the manager who will block anything that touches their team, the integration everyone assumes but nobody documented: none of that is in the files.

Filler. Asked to produce a complete-looking document, a general assistant tends to pad empty sections with generic text ("ensure a seamless user experience") that sounds fine and says nothing.

Over-confident tone. A draft written in fluent, certain prose is harder to question than a rough note. The polish hides which parts came from a source and which were inferred.

Mixed-up sources. When several documents disagree (an old deck says three regions, a newer call says two), a model may silently pick one or blend them.

A verification routine

A short routine catches most of these problems. Run it before anything goes to a client.

  1. Trace every number and name. For each metric, date, figure and person in the draft, find where it came from. If you cannot, delete it or mark it as an open question.
  2. Ask for quotes. Ask the model to back each requirement with the sentence from the source that supports it. Anthropic's guidance on reducing hallucinations recommends exactly this: extract word-for-word quotes first, and retract any claim that has no supporting quote.
  3. Allow "I don't know". Tell the model that an empty section is better than a guessed one. The same guidance notes that giving explicit permission to admit uncertainty reduces false information.
  4. Check the empty sections. An empty section is useful information: it tells you what to ask in the next interview.
  5. Look for conflicts between sources. Where two documents disagree, write down both versions and ask the client which is current.
  6. Read it as the client would. Would the decision maker recognise every statement as true? Anything that would surprise them needs a source or a question.
  7. Keep a record of changes. Know which parts were AI-drafted, which a person edited, and what changed after each new input.

Inputs matter: decks, transcripts and notes

The draft can only be as good as its inputs. Different sources carry different kinds of truth.

Input Good for Watch for
Sales or pitch decks Background, goals, the client's own framing Aspirations stated as facts; old numbers
Call and workshop transcripts Real needs, pain points, priorities in the client's words Off-hand ideas taken as requirements; who said what
Existing specs and tickets Concrete requirements and constraints Out-of-date decisions that were later reversed
Spreadsheets Figures, lists of users, data volumes Formulas and context lost when read as text
Diagrams and screenshots Current flows and systems Partial pictures that look complete
Your own notes Context nobody else wrote down Your assumptions presented as the client's

Three habits improve the result more than any prompt: give the model the newest version of each document, remove material that is clearly obsolete, and add a short note of your own with what you know that is not in the files.

Also check what happens to the files. Discovery material often contains a client's plans, names and figures, so find out where a tool sends it, whether it is stored, and whether it is used for training.

What stays a human decision

AI can propose; people decide. These parts of discovery should never be left to a model:

Decisions are worth recording with the options that were compared, so the reasoning survives after the meeting. The same example report keeps them as numbered decision records next to the diagrams they affect:

The Architecture section of the example report for the invented Fieldwise project: three diagrams with the requirements they cover, and three decision records with status and the number of options compared

Choosing a tool: questions to ask

A general assistant such as ChatGPT can do much of the drafting if you paste the material in and follow the routine above. A tool built for discovery should save you some of that routine. Either way, these questions help:

  1. Does it say when something is missing, or does it fill every section?
  2. Can you see and approve each proposed change before it is saved, or does it rewrite the document?
  3. Does it keep the result structured (requirements with priorities and acceptance criteria, goals with metrics), or produce one long text?
  4. Does it check consistency across sections when something changes?
  5. What happens to your files: are they stored, who processes them, are they used for training?
  6. Can text inside a document steer the AI? A file can contain instructions; a careful tool treats file content as data.
  7. Can you export the result in the formats your client reads?

As one example of how a dedicated tool answers these: Discovery Phase AI reads uploaded files in memory without storing them, is instructed to leave a section empty rather than invent facts, names or numbers, and shows each drafted item with a checkbox so nothing is saved until you choose it. Every later AI suggestion is a proposal you apply or reject. Its security page describes how files and AI requests are handled, including that the AI provider does not train on the content under its commercial terms. None of that removes the need for the routine above; it just makes it shorter.