The Transcript Is the Easy Part. The Certificate Is the Product.
Speech recognition now produces text of a proceeding faster and cheaper than a human can. It does not produce a record. What makes a transcript usable in a proceeding is a signed personal attestation containing four claims — and an automated pipeline cannot truthfully make a single one of them.
Two regimes, and teams see neither until a filing deadline. The first is licensure: court reporting is a licensed profession in many states, with unauthorised-practice provisions, and courts and agencies set rules for how their own records are made. The second is evidentiary: even where no licence is required, the transcript has to survive a challenge from the party it hurts. A product can be excellent at recognition and fail both.
Four Links, and Automation Only Owns Two
A transcript of record is the end of a chain. Each link has a requirement written for humans, a place where automation genuinely helps, and a specific way the automated version breaks it.
Capture
- What the record requires
- A complete, continuous recording of the proceeding, with the capture device, its placement and its operator identified, and with gaps and interruptions noted on the record as they occur.
- Where automation helps
- Automation is genuinely good here — multi-channel capture, per-speaker channels and hardware redundancy beat a single microphone.
- How it breaks
- Automatic silence trimming, voice-activity gating and lossy compression tuned for conferencing all discard audio. A record with inaudible passages is a defect; a record with passages the system decided were not speech is a defect nobody logged.
Attribution
- What the record requires
- Every utterance assigned to the correct named participant, with roles identified, and with the reporter's own knowledge of who was in the room supporting the assignment.
- Where automation helps
- Diarisation clusters voices and can label them from an attendee list.
- How it breaks
- Clustering is not identification. Overlapping speech, similar voices, remote participants on one shared line and mid-proceeding seat changes produce confident mislabels — and a misattributed admission is a substantive error, not a typo.
Transcription
- What the record requires
- A verbatim rendering, including the false starts, interruptions and objections that a summary would remove, with unintelligible passages marked as such rather than guessed.
- Where automation helps
- Recognition accuracy on clean audio is high and improving.
- How it breaks
- Language models complete rather than mark. The failure mode is not a blank where the audio was bad — it is a fluent, plausible sentence nobody said, inserted at exactly the moments the audio was hardest, which are disproportionately the contested ones.
Certification
- What the record requires
- A signed statement by a person, attesting that they were present or personally supervised the capture, that the transcript is a true and correct record of the proceeding, that they have no interest in the outcome, and identifying the oath administered.
- Where automation helps
- None. Every element is a personal assertion about a human being's own knowledge and independence.
- How it breaks
- This is the link that cannot be automated at all, and it is the one product teams discover last — usually when a client asks for a transcript that can be filed.
The Certificate, Clause by Clause
Read what the reporter is actually signing. Each clause is a statement about a person's own perception and independence, which is why no amount of model quality reaches it.
"I was present at the proceeding"
What software has: A device was present. Presence in the certificate means a person who could perceive the room, notice a speaker who was cut off, and know that the person who said the words was the person named.
"The witness was duly sworn"
What software has: Administering an oath is an act by an authorised officer. A pipeline can transcribe the words of an oath; it cannot administer one, and it cannot attest that the person raising their hand was the deponent.
"This is a true and correct transcript"
What software has: A confidence score is a statistical property of an output, not a truth claim about an event. No score converts into the assertion the certificate makes, and the passages with the lowest scores are the ones the reader most needs to trust.
"I am not related to or interested in the outcome"
What software has: Independence is a fact about a person. It is also a fact about your company when your customer is one party to the dispute and the transcript will be offered against the other.
Six Settings, Six Different Regimes
"Legal transcription" is not one market. The rules governing the record change completely between these, and a product sold across all six inherits the strictest one in the mind of every customer who gets it wrong.
Testimony under oath for use in court. States differ on whether the record must be made by a licensed reporter or may be made by other means with safeguards, and several require an officer authorised to administer the oath to be present. A software pipeline satisfies neither role.
The court decides how its own record is made. Some jurisdictions permit digital recording with a trained operator and a transcriptionist; the entitlement to appeal on a complete record is what the rules are protecting, and an unreviewable ASR output does not carry it.
Parties can generally agree on the method. The exposure is downstream: a challenge to the award, or an application to confirm it, may need a record whose accuracy someone will attest to under penalty. Agreement on convenience is not agreement on evidentiary weight.
Agencies frequently specify how the hearing record is produced and preserved, sometimes naming a reporter class. The appeal is on the record, so a defect is not a service issue — it is a due-process issue for the party who lost.
Nothing requires a licensed reporter, and the transcript may still end up as a central exhibit. Treat accuracy and custody as if it will be challenged, because the ones that matter are the ones that are.
This is the safe core of the market. Recording-consent law still applies in full, but the certified-record analysis does not — which is exactly why products should be explicit that this is what they are, rather than implying more.
The Interpreter Layer Is a Second Licence
If your product also translates, you have crossed a different certification perimeter, and the substantive failure is worse than the status failure.
| Dimension | What the rule requires | What the model does |
|---|---|---|
| Who may interpret in a proceeding | Court and agency rules commonly require a certified or otherwise court-approved interpreter, with a separate oath to interpret faithfully. | Machine translation performs the function without the certification or the oath, and there is no mechanism to qualify it as an interpreter. |
| What faithfulness means | Interpretation must preserve register, hesitation and ambiguity, not just meaning — the witness who hedges must hedge in the record. | Translation systems normalise. Fluent output smooths exactly the hesitation and equivocation that the questioner was probing. |
| Handling of the untranslatable | An interpreter is expected to flag terms with no equivalent and to seek clarification on the record. | A model produces the nearest plausible equivalent silently, and the record shows no ambiguity ever existed. |
| Language access obligations | Public entities and federally funded programmes have their own language-access duties, independent of court rules. | Deploying machine translation to satisfy a language-access obligation can fail it in both directions — inaccurate output for the individual, and a documented failure for the entity. |
The lowest-confidence-segment audit
Take any transcript your system has produced. Sort the segments by confidence ascending and read the worst twenty against the audio.
If any of them reads as a clean, grammatical sentence rather than an unintelligible marker, your pipeline is inventing testimony at precisely the moments the parties are going to fight about — and it is doing so with no visible signal to the reader.
Frequently Asked Questions
Can AI transcription replace a court reporter?
It can replace part of the work and not the part that gives a transcript its legal status. A transcript intended for use in a proceeding derives its force from a certificate, and that certificate is a set of personal attestations: that the certifier was present or personally supervised the capture, that the witness was sworn by an authorised officer, that the transcript is true and correct, and that the certifier has no interest in the outcome. Software can perform capture and produce text; it cannot make any of those four statements. Separately, court reporting is a licensed profession in many states, with unauthorised-practice provisions attached, and several jurisdictions specify by rule how the record of a proceeding may be made. The realistic posture for a product is to do the capture, diarisation and draft transcription well, and to route certification through a licensed or authorised human who reviews against the audio — not to market the output as a transcript that can be filed.
Is a transcript produced by AI admissible?
Admissibility and certification are different questions, and conflating them is the usual mistake. A recording and a transcript of it can often be admitted with a witness who can authenticate them — someone who can say what the recording is, that it fairly represents what happened, and how it was kept. That is available to an AI-generated transcript if a person with knowledge will testify to it. What is not available is the shortcut: a certified transcript is self-authenticating in practice and is accepted without a live authenticating witness, and that shortcut is exactly what the certificate buys. The second issue is challenge. An opposing party who disputes a passage can force a comparison against the audio, and an ASR pipeline that has no verbatim discipline, no marking of unintelligible passages and no retained per-segment provenance will lose that comparison — and one demonstrated fabrication contaminates the reader's confidence in the entire document.
What is the single most dangerous failure mode of ASR in a legal record?
Fluent fabrication at low-confidence moments. A traditional recording failure is obvious — a blank, a burst of noise, an inaudible marker — and everyone downstream knows to be careful there. A language-model-assisted transcription pipeline does the opposite: where the acoustic evidence is weakest, it produces its most probable completion, in grammatical English, with no visible marker. The passages with the worst audio are systematically the contested ones: people talking over each other during an objection, a witness dropping their voice on an admission, a remote participant on a bad connection. So the errors are not randomly distributed across the document; they cluster precisely where the document matters. Two controls address this specifically. Never allow the pipeline to complete over a low-confidence span — emit an explicit unintelligible marker instead. And retain per-segment confidence and the aligned audio offset, so any disputed line can be checked in seconds rather than argued about.
Do we need consent to record a proceeding or an interview?
Almost always, and it is a separate regime from everything else in this article. Recording-consent law varies by state, with some requiring all parties to consent, and it applies to the audio your product captures regardless of what you later do with it. Multi-jurisdiction participants on a remote call raise the question of which state's rule governs, and the conservative answer is the most protective one present. Courts and agencies add their own layer: many prohibit recording a proceeding without permission entirely, including by a participant, and a notetaker that joins a remote hearing can violate that rule without anyone in your company knowing the hearing was a hearing. If your product auto-joins calendared meetings, it will eventually auto-join a legal proceeding. Build the ability to detect and block that, and make the announcement of recording explicit and audible rather than a line in a settings page.
How does machine translation interact with interpreter certification?
It sits on the wrong side of a separate licensing line. Where a proceeding involves a participant with limited proficiency in the language of the proceeding, court and agency rules commonly require a certified or court-approved interpreter who takes a separate oath to interpret faithfully. Machine translation performs the function and cannot take the oath or hold the certification, and there is no procedure to qualify software as an interpreter. Beyond status, there is a substantive problem: interpretation must preserve register, hesitation and ambiguity, and translation systems are built to normalise them away. A witness who hedges in one language becomes a witness who states plainly in another, and the hedge was often the point of the question. Machine translation has real uses in this space — preparation, review, search across a multilingual corpus — but placing it in the interpreting seat of a proceeding creates a defect in the record and, depending on the setting, a language-access failure on top.
We sell to law firms but we are not in the courtroom. Does any of this reach us?
It reaches you through what your customers do with the output and through what your marketing implies they can. Two exposures matter. First, claim substantiation: describing output as a certified, court-ready or verbatim transcript when it is an uncertified machine draft is a straightforward deceptive-claims problem, and it is the sort of claim that surfaces in a filing rather than in a support ticket. Second, professional-responsibility spillover: a lawyer who relies on an inaccurate transcript has their own competence and candour duties, and vendors get named in the resulting mess. There is also a confidentiality layer that is easy to miss — proceeding audio contains privileged and often sealed material, so retention windows, subprocessor lists, training use and access logging need to be tighter than a general transcription product's defaults, and a firm's outside-counsel guidelines will usually say so before your security questionnaire does.
What does a defensible AI transcription workflow for legal use look like?
Five properties, in order of how often they are missing. One: unmodified source audio retained with a hash, so any later dispute compares against the original rather than a processed derivative, and with no silence trimming or aggressive noise suppression applied to the retained copy. Two: per-segment confidence and audio offsets retained alongside the text, so a challenged line is verifiable rather than defended by assertion. Three: explicit unintelligible markers, with the pipeline prohibited from completing across a low-confidence span. Four: a named human reviewer who checks against the audio, with the review recorded — who, when, which segments changed — and certification issued by a person authorised in that jurisdiction, never by the system. Five: honest product language that distinguishes a working draft from a record, because most of the legal risk in this category arrives through the marketing page rather than the model.
Related Reading
- AI meeting notetakers and recording consent — the consent regime that applies to the audio before any of this does.
- AI tools, privilege and work-product waiver — what happens to privileged material that passes through a vendor.
- Speech recognition and accent discrimination — whose words the pipeline gets wrong, and why it is not evenly distributed.