Scribe meeting transcription: extract action items with quotes and timestamps

Turn a recording into a Scribe transcript and proposed action items with exact quotes and timestamps. Follow a real API run using a fictional meeting sample.

Rubber stamp and ink pad on a pale card, dusty-pink background, olive dots and the title Scribe: Meeting Notes.

A useful meeting workflow has three separate outputs: the recording, a transcript that preserves what was said, and proposed action items linked to the relevant words and times. A fluent summary is not a substitute for the transcript. An item in a generated JSON file is not an approved task in your project system.

This tutorial uses Ofox to transcribe a 43.92-second synthetic recording with elevenlabs/scribe_v2, then passes the timestamped response to openai/gpt-6-luna. The recording deliberately includes an assigned task, an unassigned task, an undecided release date and an unapproved budget. The text model returned proposed items and open questions without sending messages or creating external tasks. The source files and raw responses are downloadable.

The meeting is fictional, written for this tutorial and voiced synthetically. Names are part of the script, not real participants or verified speaker identities. The example demonstrates an auditable processing chain; it is not evidence of performance on noisy human meetings or reliable automatic diarization.

Define a deliverable that can be reviewed

Before uploading audio, decide what a reviewer should receive. For this workflow, the minimum bundle contains the original recording, raw ASR response, any corrected transcript, proposed actions with quotations and time ranges, unresolved questions, and a review decision. Keep these artifacts separate so that later corrections do not erase what the models actually returned.

An action record needs more than an attractive sentence. It should distinguish the task from its owner, deadline and approval status. If a field is not supported by the recording, use null or a visible “needs confirmation” state. Completing every table cell by guessing makes the document less useful, not more finished.

The working kit includes meeting-script.txt, meeting audio, meeting-transcript.json, extract_actions.py, the extraction request and raw response. These files let you verify the chain without uploading someone else’s private meeting.

Fictional meeting audio

1. Obtain authorized input and preserve its timeline

Use a recording you are permitted to process with the selected cloud service. For a workplace meeting, follow the organization’s recording, retention and access rules. Do not include passwords, private access tokens or unrelated personal information merely because the transcript might help the model understand a task.

Keep the original file and create a supported audio copy if needed. The Ofox route tested here accepts WAV or MP3. If you extract an audio track from video, record which track was selected and whether any leading time was trimmed. Every later quotation must refer to the correct timeline.

For this example, we generated one 43.92-second MP3 from the original script. It includes these deliberately different kinds of statements:

  • Leo will check the API example by October 15, 2026.
  • Nobody has been assigned to record the Japanese narration.
  • The speaker cannot approve the production release.
  • The release date is undecided, so invitations must not be sent yet.
  • A thirty-dollar test budget was discussed but not approved.

These contrasts are the test’s point. A system that extracts “send invitations” while losing “do not” can produce a plausible-looking but harmful work list. A system that treats discussion of a budget as authorization has crossed from summarization into invention.

2. Transcribe with Scribe and retain the original response

Set OFOX_API_KEY, install requests, FFmpeg and ffprobe, then run the provided client from the kit directory:

python3 audio_api.py transcribe \
  --input meeting.mp3 \
  --output my-meeting-transcript.json

The client uploads multipart form data to /v1/audio/transcriptions using elevenlabs/scribe_v2 and verbose_json. Check the Ofox Scribe model page for the current exposed route. This example does not assume every parameter in the native provider API is accepted by the gateway.

Save the response before editing it. The sample JSON contains full text plus word-level start and end times. It also normalizes the spoken date into “October 15th, 2026.” That written normalization is useful, but it is not evidence that the model knows the current date or that a relative deadline can always be resolved safely.

The recorded audio ends at 43.92 seconds, and the last returned word ends at 43.74 seconds. Keep duration and usage fields distinct. A final spoken-word timestamp does not replace file duration or settle the amount charged for a transcription request.

If the API fails, preserve the error body and request ID rather than passing the error message to the extraction model as a transcript. If it returns text without timings, you can still draft a summary, but you cannot claim to have verified time-linked evidence until alignment is available.

3. Review the transcript where mistakes change decisions

Review names, dates, numbers, negations and permission statements first. In this synthetic example the source script is available for comparison. In a real meeting, the recording remains the primary evidence; slides, agendas and earlier notes may clarify vocabulary but should not silently overwrite what a participant said.

Keep an edit log for corrections. A row can contain the original ASR phrase, corrected phrase, time range, reason and reviewer. For unresolved audio, mark uncertainty. Do not invent a named owner merely because a nearby person was speaking or because that person usually handles the work.

This response does not provide verified speaker labels. “Leo speaking” appears in the deliberately spoken script. It is text in an audio file, not acoustic proof that a particular employee spoke. The extraction prompt explicitly warns the next model not to infer speaker identity from that convention.

When multiple people interrupt or disagree, a clean action list may require additional review. The short, single-voice tutorial recording does not test overlap, background noise, cross-talk or accents. Treat those as separate evaluation conditions before using an automated workflow on your team’s recordings.

4. Ask the text model for evidence-linked proposals

The extraction stage receives both the text and the actual word array. It is instructed to use only that material, return unknown fields explicitly, keep negations and distinguish commitments from suggestions. It must not send messages or modify a project system.

The complete request is in actions-request.json. Its core instruction is:

Use only the supplied ASR words and text.
Return actions, decisions, and open_questions.
Each action needs task, owner, due_date, status=proposed,
an exact supporting quote, and start/end seconds copied
from the word timestamps. Use null for unknown fields.
Separate commitments, proposals, and negations.
Do not infer acoustic speaker identity from a spoken name.
Do not invent approvals, send messages, or create tasks.

Run extract_actions.py to reproduce the included sample extraction. It uses openai/gpt-6-luna through /v1/chat/completions, reads the timestamped transcript and saves the raw result. The shipped script refuses to overwrite an existing actions-response.json; choose a separate working copy or versioned output before intentionally making another paid request.

The example requests JSON through its instructions and then preserves the response for validation. It does not claim that a structured-output schema was enforced at the API level. In a production system, validate the parsed object against your own schema even if the model’s message looks like valid JSON.

For a new meeting, change the input file deliberately and preserve the prompt version. Do not casually reuse a script that still points to the tutorial’s transcript. A perfectly formatted result from the wrong recording is still the wrong result.

5. Examine the actual result, including its imperfections

The sample model returned a nested actions object with commitments, proposals and negations, plus decisions and open_questions. All action statuses are proposed. The raw response is evidence of one output; it should not be treated as a stable API schema across future calls.

Returned itemSupporting intervalReview interpretation
Check the API example; owner Leo; due 2026-10-1511.18–15.64 sAn explicit assignment in the transcript, still a proposed task record
Prepare the internal product demo7.78–10.58 sGoal identified, owner and deadline remain unknown
Japanese narration owner16.26–19.26 sOpen question, not an assigned task
Release date27.28–29.24 sUnresolved; do not derive a deadline from the API review date
Do not send invitations yet29.66–32.14 sConstraint that must survive summarization
Discussed test budget32.74–36.66 sNo spending authorization

The raw response also contains “I can review the example” as a separate proposal. That can overlap with the API-check assignment. A reviewer should decide whether these refer to the same work before exporting tasks. Automatically creating both could produce duplicate assignments even though each quotation is real.

The model left decisions empty and preserved the unapproved budget as an open question. That is preferable to inventing approval, but it is still only a result for this carefully constructed example. A system needs broader tests before relying on this behavior across real discussions.

A useful review outcome might retain the explicit API check, keep the demo as an unassigned proposal, retain the three open questions, and store the invitation restriction as a constraint rather than a task to execute. Keep that editorial interpretation separate from actions-response.json so the original output remains inspectable.

6. Validate quotations, timestamps and missing fields

Start with deterministic checks. Parse the JSON, verify required keys, and reject malformed times. Every cited interval should fit inside the recording and should correspond to the beginning and end of the quoted words. The fact that a number falls within the file duration is necessary but not sufficient evidence that it points to the right sentence.

Check exact quotes against the transcript. If you normalize whitespace or punctuation for matching, record that rule. Do not use a broad fuzzy match that could treat “approved” and “not approved” as interchangeable. For the sample, the budget quotation includes the denial of approval; dropping that clause would change the meaning.

Then perform semantic review. A valid quote can still support the wrong conclusion. “I cannot approve the production release” supports a limitation, not permission to deploy. “The release date is still undecided” cannot support a due date inferred from another sentence’s API-review deadline.

Unknown values should remain visible in the interface. A blank owner should not silently become the uploader, meeting organizer or first named person. A missing deadline should not default to today. If a downstream tool requires those fields, keep the item in a review queue rather than manufacturing data to satisfy the form.

7. Keep extraction separate from task creation

The tutorial stops at proposed records. It does not create a ticket, send an email or schedule a calendar event. That boundary is important because transcription and interpretation can be wrong even when both API calls return HTTP 200.

Before connecting a project tool, define who can approve an action and what evidence they see. Show the quote and audio interval beside the proposed owner and deadline. Record acceptance, correction or rejection, and maintain an audit trail linking the exported task to the reviewed version.

Use an idempotency or deduplication key for eventual task creation, based on a reviewed meeting version and action identifier. Re-running summarization should not create duplicate tickets. A revised action should update the intended record only after its changes have been reviewed; it should not silently overwrite an earlier approved commitment.

Meeting participants may also change their minds after the recording. A later human decision belongs in the task history, with its own source, rather than being retroactively inserted into the transcript as if it had been spoken at the meeting.

8. Troubleshoot the workflow one stage at a time

If a transcription request fails, check supported input formats, credentials and the exact error. Earlier Scribe calls for this project hit an upstream quota issue; restoration was confirmed by new successful requests. Wallet status alone was not enough to diagnose the route.

If transcription is incomplete, inspect the input audio, its selected track and whether the upload finished. If the text model returns malformed JSON, keep the raw response and fix the extraction or validation process. Do not repeatedly upload the same audio when only the downstream JSON formatting failed.

If the output invents owners or dates, strengthen the evidence requirement and review actual examples. A longer prompt does not guarantee compliance. Keep test cases containing unassigned work, tentative dates, denied permissions and contradictory statements, and require the validator and reviewer to catch them.

For very long meetings, preserve chunk offsets and reconcile repeated or superseded decisions across sections. The small script in this kit is a transparent short-recording example, not a complete long-meeting orchestration system. Do not advertise a long-form limit or automatic cross-meeting memory that has not been tested.

9. Measure useful output rather than summary length

Track how many proposed items were accepted unchanged, corrected, rejected or merged as duplicates. Record time spent reviewing and the number of missed commitments found by a reviewer. Those measures are more informative than the number of generated bullet points.

Keep transcription and text-model usage separate when checking costs. A per-second ASR quantity and a text-token quantity are different units. The artifact metadata is useful for matching requests to your own billing records, but this tutorial does not present an unreconciled estimate as an actual invoice.

For publication or team deployment, the acceptance condition is a reviewer-approved, source-linked work list with unresolved fields preserved. A polished summary that loses the invitation restriction or fabricates budget approval fails that condition, even if it is short and easy to read.

Frequently Asked Questions

Is this a recording of a real customer meeting?
No. It is an original fictional script voiced synthetically for the tutorial. The API transcription and text-model extraction are real; the meeting and participants are not.
Does Scribe identify the real speakers in this example?
The saved response does not provide verified speaker identities. Spoken labels in the script are not evidence of acoustic identity, so the workflow avoids inventing attribution.
Why are all action statuses proposed?
Extraction is not approval. A reviewer should confirm the owner, deadline, evidence and permission before creating or executing a task.
Can I automatically send the generated tasks to my team?
This example does not do that. Add an explicit review and deduplication stage before any external task creation or notification, and preserve the supporting source for each approved item.