Data type
Text and subtitle import
Import TXT, Markdown, SRT, and VTT through the main application or API with versioned text/time locators.
Implemented task
Public task
Open source actions and choose Add source.
Expected outcome: A validated source enters the durable revision and processing pipeline.
Declared product contract
Inputs and task boundary
Inputs
- UTF-8 TXT
- UTF-8 Markdown
- SRT subtitles
- VTT subtitles
Input constraints
- Direct text files must be UTF-8 and no larger than 50 MB.
User actions
- Open Add source in a project workspace.
- Choose a text or subtitle file and review detected structure.
- Submit the source to create a versioned transcript.
Outputs
- A durable source with versioned text and time locators where the format supplies time.
Decision boundary
Where it fits
- Prepared transcripts, notes, and subtitle files in the declared formats.
Outside the boundary
Where it does not fit
- Non-UTF-8 text or files above the declared size limit.
Current scope
Limitations
- Direct text files must be UTF-8 and no larger than 50 MB.
Verification
Product proof
journey
proof.m1-material-import
- Verified
- Review due
- Expires
- Eight public fixture formats enter the real source and revision pipeline.
- Exact locators and manual coding persist across refresh.
- Withdrawal previews impact and erases live database and stored objects.
Failure boundary: The proof does not claim M3 typed cases and variables or frame-level video analysis.
journey
proof.m10-transcript-source-handoff
- Verified
- Review due
- Expires
- A local transcript inspection remains browser-only until the user selects a project and requests the official source preview.
- The exact normalized source preview renders intake fields, hash, changes, and a virtualized 40-segment transcript while PostgreSQL and storage contain zero new source objects.
- Only a digest-bound final confirmation creates one source, one pipeline run, one queued pipeline job, one stored file, and one source-uploaded audit event.
- The source acquisition lineage and audit preserve the transcript inspector tool ID, version, result ID, input fingerprint, and free-tool acquisition context including sample and template lineage.
- The successful import retains the browser draft until the user explicitly deletes it.
Failure boundary: This journey proves the transcript-inspector source adapter in writable single-user mode, including one sample-origin and template-origin acquisition lineage path; other source handoff tools and multi-user role recovery require their own journeys.
external standard
proof.m10-open-standard-formats
- Verified
- Review due
- Expires
- A claim-driven inventory applies independent standard checks only to implemented QDC/QDPX-codebook, WebVTT, DOCX, and XLSX claims.
- Both QDC and the QDPX codebook projection validate against the vendored official REFI-QDA Codebook 1.0 schema.
- WebVTT and OOXML fixtures pass independent parsers that do not call the product import parsers.
- Malformed QDC, WebVTT, and DOCX failure injections fail closed.
- Complete-project QDPX, saved queries, report builder, and all partial/planned capabilities remain explicitly unproved.
Failure boundary: This proof covers implemented open-format boundaries only; it does not certify complete-project QDPX, third-party round trips, saved queries, report builder behavior, or a production service level.
OpenVerbatim is an open-source (Apache-2.0) qualitative data analysis platform for coding and analyzing interview transcripts. AI-suggested codes stay marked as suggestions until a human reviewer confirms or rejects them, and every decision is kept in an audit trail. The full feature set is available when self-hosted; there is no paid feature wall.