WeXL Forge · Specify · Beta

Forge Spec

Turns notes, transcripts and docs into epics, stories and acceptance criteria

Takes
Meeting transcript (.txt, .md, .vtt)Notes or a requirements doc (.md, .docx, .pdf)
Gives
Epics and user storiesGiven/When/Then acceptance criteriaOpen questions and assumptionsMarkdown or Jira CSV
Runs on
Claude CodeKiroOpenCode
Updated

Beta: in use on our own product work; not yet handed over to a client team.

The problem

Most requirements start life as a conversation. A product manager, a tech lead, a designer and a tester spend twenty minutes on a call and agree a dozen things: what the feature does, what it must never do, the edge cases, what is parked for later. Then someone has to turn that call into stories a team can size and build. That job usually lands on whoever is least busy, a day or two later, from memory and a few scribbled notes.

That is where requirements quietly break. A decision made at minute nine gets lost. An idea someone floated and the group rejected turns up as a story. “No timer on retries” never becomes an acceptance criterion, so nobody tests for it. The stories that reach the sprint are tidy, but they no longer match what was agreed, and nobody can tell which lines came from the call and which the writer filled in.

Good looks like this: every story traces to something that was actually said; every acceptance criterion can be tested as written; what was parked is listed as parked; and the questions nobody answered are at the top, where the tech lead will see them before sizing.

How it works

Forge Spec reads whatever the conversation left behind: a transcript with timestamps, meeting notes, a requirements document, or several of these together. It works in three passes.

First, it extracts: decisions, constraints, edge cases, ideas that were raised and parked, and questions left open. Each item keeps a reference to where it came from, a timestamp in a transcript or a section in a document. A decision that was reversed later in the call is recorded as the later decision.

Second, it writes: an epic with a one-line goal, then user stories, each with Given/When/Then acceptance criteria. Every criterion carries the reference it was built from. Anything the stories depend on that nobody said goes into a separate Assumptions list, never into the stories themselves.

Third, it checks its own output against your team’s rules: your story template, your definition of ready, your naming, and your glossary of product terms. These live in the steering and skills files that come with the Workbench, so the output reads like your team wrote it, not like a generic template.

A person then reviews it. In practice that is the product owner or business analyst, and the review is quick because the open questions and assumptions are listed first and every line has a source to check against. Forge Spec writes Markdown by default, or a Jira-ready CSV once the review is done.

It runs inside the tools your team already uses: Claude Code, Kiro or OpenCode, with the model your rules allow.

Where it gets it wrong

Forge Spec is good at structure and traceability. It is not a substitute for the people in the room, and it fails in predictable ways.

  • It can treat an idea as a decision. When someone says “what if we also…” and the call moves on without a clear no, the agent may not know the idea was dropped. We mitigate this by recording only explicit agreements as decisions; anything unclear goes to open questions, and parked ideas are listed as out of scope so a reviewer can see them.
  • It cannot know what was never said. Non-functional needs such as performance, accessibility, audit or data retention are often assumed by the team and never spoken. Forge Spec flags categories that are missing, but it cannot invent the right answer, and it should not.
  • Story size is a judgement call. It sometimes splits a feature finer than your team would, or bundles two things your team keeps apart. Your definition of ready and a few examples of good stories in the steering files reduce this; the reviewer still has the final say.
  • Transcripts are imperfect. Speaker labels get swapped, product names get misheard, and crosstalk loses words. A glossary of product terms fixes most misheard names. When the speaker matters, for example who owns an action, the agent quotes the line rather than guessing.

We would rather the agent leave a question open than answer it confidently and wrongly. The review step exists because of this section.

Why not just prompt a chatbot

You can paste a transcript into a chatbot and get user stories back in seconds. They will look plausible, and that is the problem: you cannot tell which lines came from the call and which were filled in. What makes Forge Spec useful is not the model. It is the steering that encodes your template and definition of ready, the skills that force a source on every criterion and keep assumptions separate, the checks that run before a person sees the output, and the fit to your tools and tracker. That is the part we build, and the part you keep.

What you get

  • In your engagement: Forge Spec runs on your refinement calls from week one, and your product owners review its output instead of writing stories from scratch.
  • Under your brand: services firms can offer it inside their own delivery platform, with their story standards in the steering files.
  • Built to own: we tune it to your template, glossary and tracker, and hand over the steering, skills and source.

See three ways to use an agent.

Worked example

An illustrative refinement call about a FluentEdge feature, written for this page; the people are named by role. The output shows exactly the format Forge Spec produces.

InputRefinement call transcript, 21 minutes
Raw
Refinement call: "Retry what I got wrong" for FluentEdge Practice BETs
Attendees: Product manager (PM), Tech lead (TL), Designer (DS), QA lead (QA)
Duration: 21 minutes. Transcript lightly cleaned of filler words.

[00:00:15] PM: Okay, one item today. Learners finish a Practice BET, see their score, and then the only option is to retake the whole thing. Eight questions. Support tickets say people want to redo just the ones they got wrong.
[00:00:41] TL: Is this all four skills or just reading and listening? Speaking and writing are AI-evaluated, there isn't really a "wrong" answer.
[00:00:58] PM: Good point. Let's say for now, reading and listening only. Those are the MCQ ones.
[00:01:20] DS: Where does the button live? On the result screen, I assume. Next to "Retake test".
[00:01:31] PM: Yes. Something like "Retry the 3 you missed". With the count.
[00:01:44] QA: What if they got everything right?
[00:01:49] PM: Then don't show the button. Maybe a "Nice work" line instead.
[00:02:05] TL: Does a retry count towards their score? Because the Summative score is what recruiters see.
[00:02:16] PM: No. Retries are practice only. They should not change the Practice BET score that's already recorded, and definitely nothing on the Summative.
[00:02:40] DS: Should the retry show the same options in the same order? If it's the same order they'll just remember the position.
[00:02:52] TL: We can shuffle the options. The question bank already supports that.
[00:03:05] PM: Shuffle, yes.
[00:03:30] QA: Do we show the correct answer after the retry, or during?
[00:03:41] PM: After. Same as the normal test. Show which ones they now got right, and for anything still wrong, show the correct answer and the explanation.
[00:04:10] TL: The explanations only exist for about 70 percent of the reading questions. Listening has almost none.
[00:04:22] PM: Then show the correct answer, and the explanation where we have it. Content team can backfill. Let's not block on that.
[00:05:02] DS: Can they retry the retry? Like keep going until they get everything?
[00:05:12] PM: I'd say yes, until they've got them all right. Or they can stop any time.
[00:05:25] TL: That's a loop. Do we need a cap? Each attempt writes rows.
[00:05:36] PM: Hmm. I don't know. Let's not cap it for now but I want to know how many attempts people actually do.
[00:05:50] TL: So an event per retry attempt. I'll make sure analytics gets attempt number and how many were still wrong.
[00:06:30] DS: One idea: what if we also let them retry the questions they got right but were slow on? Like over 60 seconds.
[00:06:44] PM: Interesting, but no, not in this one. Keep it to wrong answers. Park it.
[00:07:15] QA: Mobile. The result screen on mobile already has two buttons stacked. A third makes it long.
[00:07:28] DS: I'll make "Retry what you missed" the primary on both. "Retake test" becomes a text link. That's actually the better default anyway.
[00:07:40] PM: Agreed.
[00:08:20] TL: What about timing? The real Practice BET is timed. Is the retry timed?
[00:08:31] PM: No timer on retries. It's for learning, not for measuring.
[00:09:02] QA: Is this for all users or only the paid tier?
[00:09:10] PM: I need to check with the business side. Write it down as open.
[00:09:45] TL: Offline? Some institutions use it on patchy wifi.
[00:09:55] PM: Same behaviour as the normal test. If the normal test needs a connection, so does this.
[00:10:30] QA: And if they close the app in the middle of a retry?
[00:10:39] PM: They lose that retry, no big deal, they can start again from the result screen. Don't build resume.
[00:11:15] TL: One more. Institutions see reports. Do retries show in the teacher dashboard?
[00:11:27] PM: Not in this release. But the data should be there so we can add it later.
[00:12:00] DS: Copy for the button. "Retry the 3 you missed" or "Practise the 3 you missed"?
[00:12:12] PM: "Retry the 3 you missed". Plain.
[00:12:30] QA: Singular. "Retry the 1 you missed" reads odd.
[00:12:38] DS: "Retry the one you missed" when it's one. I'll write both strings.
[00:13:05] PM: Good. Target is the sprint after next. Tech lead, size it once the stories are written.
[00:13:20] TL: Will do. Send me the stories with the open questions on top.
[End of call]
OutputEpic, five stories, open questions
Raw
# Epic: Retry the questions you missed in a Practice BET

**Goal:** Let a learner practise only the reading and listening questions they got wrong in a Practice BET, without changing any recorded score.
**Source:** Refinement call, 21 min. References in [brackets] are transcript timestamps.

## Open questions (resolve before sizing)

1. **Which plans get this?** All users, or the paid tier only? Owner: PM, with the business side. [00:09:10]
2. **Is a cap on retry attempts needed?** Agreed: no cap now; decide after seeing the attempt data from story 5. [00:05:36]
3. **Non-functional needs not discussed.** Accessibility of the new button and the retry screen, and how long retry events are kept, did not come up on the call. Confirm, or mark as not applicable.

## Stories

### 1. Offer a retry of missed questions on the result screen

As a learner who has just finished a Practice BET,
I want to retry only the reading and listening questions I got wrong,
so that I can practise my weak spots without redoing the whole test.

**Acceptance criteria**

- Given a learner got at least one reading or listening question wrong,
  when the result screen loads,
  then a primary button reads "Retry the N you missed", where N is the number of wrong reading and listening questions. [00:01:31]
- Given exactly one such question is wrong,
  then the button reads "Retry the one you missed". [00:12:38]
- Given no reading or listening question is wrong,
  then the retry button is not shown and a short "Nice work" line is shown instead. [00:01:49]
- Speaking and writing questions are never included in N or in the retry. [00:00:58]
- "Retake test" remains available as a text link below the primary button, on mobile and desktop. [00:07:28]

### 2. Run a retry set

As a learner,
I want the retry to work like a short practice round,
so that I can focus on getting the answers right rather than on the clock.

**Acceptance criteria**

- Given a learner starts a retry,
  then only the missed reading and listening questions are shown,
  and the answer options for each question are in a different random order from the original attempt. [00:02:52]
- The retry has no timer. [00:08:31]
- If the device has no connection, the retry behaves exactly as the normal Practice BET does without a connection. [00:09:55]
- If the learner leaves the app during a retry, that retry is discarded; starting again from the result screen begins a new retry. No resume. [00:10:39]

### 3. Show results and answers after a retry

As a learner who has finished a retry,
I want to see what I now got right and the correct answers for the rest,
so that I learn from the questions I still miss.

**Acceptance criteria**

- Given a learner finishes a retry,
  then each question is marked right or wrong. [00:03:41]
- For each question still wrong, the correct answer is shown, plus the explanation where one exists. A missing explanation shows the correct answer only, with no empty explanation box. [00:04:22]
- Answers are shown only after the retry is finished, never during it. [00:03:41]
- If any questions are still wrong, the learner can start another retry of just those; they can also stop at any time. [00:05:12]

### 4. Keep retries out of every score

As a learner, and as a recruiter reading a Summative score,
I want practice retries to have no effect on any recorded score,
so that scores mean the same thing they meant before this feature.

**Acceptance criteria**

- A retry never changes the recorded Practice BET score for the original attempt. [00:02:16]
- A retry never affects the Summative score or anything a recruiter sees. [00:02:16]
- Retries do not appear in the teacher dashboard in this release. [00:11:27]

### 5. Record retry attempts for analytics

As the product team,
we want each retry attempt recorded,
so that we can decide on an attempt cap and add retries to institution reports later.

**Acceptance criteria**

- Each retry attempt emits one analytics event with: learner, Practice BET attempt, attempt number, questions in the retry, and questions still wrong at the end. [00:05:50]
- The stored data is enough to show retries in the teacher dashboard later without a migration. [00:11:27]

## Assumptions (not stated on the call; confirm or correct)

- "Wrong" means answered incorrectly; an unanswered question in the original attempt also counts as missed.
- The retry uses the same question versions as the original attempt, even if the question bank has been edited since.
- Option shuffling follows the existing question-bank setting; no new content work is needed.

## Out of scope (raised and parked)

- Retrying questions answered correctly but slowly (over 60 seconds). Parked by PM. [00:06:44]
- Resuming an interrupted retry. [00:10:39]
- Showing retries in the teacher dashboard. [00:11:27]
- Backfilling missing explanations; content team, separately. [00:04:22]

Run it on your input.

Send us a real meeting transcript or set of notes. We run Forge Spec on it and send the output back within one working day, with a note on what we would change for your setup.

Start here

See Forge on your own codebase.

Start a conversationWe'll show the Live agents on a real workflow of yours.