The problem
Most requirements start life as a conversation. A product manager, a tech lead, a designer and a tester spend twenty minutes on a call and agree a dozen things: what the feature does, what it must never do, the edge cases, what is parked for later. Then someone has to turn that call into stories a team can size and build. That job usually lands on whoever is least busy, a day or two later, from memory and a few scribbled notes.
That is where requirements quietly break. A decision made at minute nine gets lost. An idea someone floated and the group rejected turns up as a story. “No timer on retries” never becomes an acceptance criterion, so nobody tests for it. The stories that reach the sprint are tidy, but they no longer match what was agreed, and nobody can tell which lines came from the call and which the writer filled in.
Good looks like this: every story traces to something that was actually said; every acceptance criterion can be tested as written; what was parked is listed as parked; and the questions nobody answered are at the top, where the tech lead will see them before sizing.
How it works
Forge Spec reads whatever the conversation left behind: a transcript with timestamps, meeting notes, a requirements document, or several of these together. It works in three passes.
First, it extracts: decisions, constraints, edge cases, ideas that were raised and parked, and questions left open. Each item keeps a reference to where it came from, a timestamp in a transcript or a section in a document. A decision that was reversed later in the call is recorded as the later decision.
Second, it writes: an epic with a one-line goal, then user stories, each with Given/When/Then acceptance criteria. Every criterion carries the reference it was built from. Anything the stories depend on that nobody said goes into a separate Assumptions list, never into the stories themselves.
Third, it checks its own output against your team’s rules: your story template, your definition of ready, your naming, and your glossary of product terms. These live in the steering and skills files that come with the Workbench, so the output reads like your team wrote it, not like a generic template.
A person then reviews it. In practice that is the product owner or business analyst, and the review is quick because the open questions and assumptions are listed first and every line has a source to check against. Forge Spec writes Markdown by default, or a Jira-ready CSV once the review is done.
It runs inside the tools your team already uses: Claude Code, Kiro or OpenCode, with the model your rules allow.
Where it gets it wrong
Forge Spec is good at structure and traceability. It is not a substitute for the people in the room, and it fails in predictable ways.
- It can treat an idea as a decision. When someone says “what if we also…” and the call moves on without a clear no, the agent may not know the idea was dropped. We mitigate this by recording only explicit agreements as decisions; anything unclear goes to open questions, and parked ideas are listed as out of scope so a reviewer can see them.
- It cannot know what was never said. Non-functional needs such as performance, accessibility, audit or data retention are often assumed by the team and never spoken. Forge Spec flags categories that are missing, but it cannot invent the right answer, and it should not.
- Story size is a judgement call. It sometimes splits a feature finer than your team would, or bundles two things your team keeps apart. Your definition of ready and a few examples of good stories in the steering files reduce this; the reviewer still has the final say.
- Transcripts are imperfect. Speaker labels get swapped, product names get misheard, and crosstalk loses words. A glossary of product terms fixes most misheard names. When the speaker matters, for example who owns an action, the agent quotes the line rather than guessing.
We would rather the agent leave a question open than answer it confidently and wrongly. The review step exists because of this section.
Why not just prompt a chatbot
You can paste a transcript into a chatbot and get user stories back in seconds. They will look plausible, and that is the problem: you cannot tell which lines came from the call and which were filled in. What makes Forge Spec useful is not the model. It is the steering that encodes your template and definition of ready, the skills that force a source on every criterion and keep assumptions separate, the checks that run before a person sees the output, and the fit to your tools and tracker. That is the part we build, and the part you keep.
What you get
- In your engagement: Forge Spec runs on your refinement calls from week one, and your product owners review its output instead of writing stories from scratch.
- Under your brand: services firms can offer it inside their own delivery platform, with their story standards in the steering files.
- Built to own: we tune it to your template, glossary and tracker, and hand over the steering, skills and source.