Data type
Document import
Import DOCX and PDF material while preserving source positions.
Implemented task
Public task
Open source actions and choose Add source.
Expected outcome: A validated source enters the durable revision and processing pipeline.
Declared product contract
Inputs and task boundary
Inputs
- DOCX document
- Text-bearing PDF
Input constraints
- PDF extraction requires Poppler.
- Scanned PDFs use the image OCR path instead of claiming embedded text.
User actions
- Open Add source and choose a document.
- Review the extracted material preflight.
- Submit it to the source and revision pipeline.
Outputs
- A durable document source with canonical text or page-position evidence locations.
Decision boundary
Where it fits
- DOCX and text-bearing PDF research material.
Outside the boundary
Where it does not fit
- Scanned PDFs treated as text without the OCR path.
Current scope
Limitations
- DOCX locations use canonical extracted-text offsets because Word pages are layout-dependent.
- PDF page regions require Poppler and scanned PDFs must be imported as images for OCR.
Verification
Product proof
journey
proof.m1-material-import
- Verified
- Review due
- Expires
- Eight public fixture formats enter the real source and revision pipeline.
- Exact locators and manual coding persist across refresh.
- Withdrawal previews impact and erases live database and stored objects.
Failure boundary: The proof does not claim M3 typed cases and variables or frame-level video analysis.
external standard
proof.m10-open-standard-formats
- Verified
- Review due
- Expires
- A claim-driven inventory applies independent standard checks only to implemented QDC/QDPX-codebook, WebVTT, DOCX, and XLSX claims.
- Both QDC and the QDPX codebook projection validate against the vendored official REFI-QDA Codebook 1.0 schema.
- WebVTT and OOXML fixtures pass independent parsers that do not call the product import parsers.
- Malformed QDC, WebVTT, and DOCX failure injections fail closed.
- Complete-project QDPX, saved queries, report builder, and all partial/planned capabilities remain explicitly unproved.
Failure boundary: This proof covers implemented open-format boundaries only; it does not certify complete-project QDPX, third-party round trips, saved queries, report builder behavior, or a production service level.
OpenVerbatim is an open-source (Apache-2.0) qualitative data analysis platform for coding and analyzing interview transcripts. AI-suggested codes stay marked as suggestions until a human reviewer confirms or rejects them, and every decision is kept in an audit trail. The full feature set is available when self-hosted; there is no paid feature wall.