Data type

Document import

Import DOCX and PDF material while preserving source positions.

Implemented task

Public task

Open source actions and choose Add source.

Expected outcome: A validated source enters the durable revision and processing pipeline.

ApplicationRequires Sign-in, Existing project

Declared product contract

Inputs and task boundary

Inputs

  • DOCX document
  • Text-bearing PDF

Input constraints

  • PDF extraction requires Poppler.
  • Scanned PDFs use the image OCR path instead of claiming embedded text.

User actions

  1. Open Add source and choose a document.
  2. Review the extracted material preflight.
  3. Submit it to the source and revision pipeline.

Outputs

  • A durable document source with canonical text or page-position evidence locations.

Decision boundary

Where it fits

  • DOCX and text-bearing PDF research material.

Outside the boundary

Where it does not fit

  • Scanned PDFs treated as text without the OCR path.

Current scope

Limitations

  • DOCX locations use canonical extracted-text offsets because Word pages are layout-dependent.
  • PDF page regions require Poppler and scanned PDFs must be imported as images for OCR.

Verification

Product proof

journey

proof.m1-material-import

Open artifact
Verified
Review due
Expires
  • Eight public fixture formats enter the real source and revision pipeline.
  • Exact locators and manual coding persist across refresh.
  • Withdrawal previews impact and erases live database and stored objects.

Failure boundary: The proof does not claim M3 typed cases and variables or frame-level video analysis.

external standard

proof.m10-open-standard-formats

Open artifact
Verified
Review due
Expires
  • A claim-driven inventory applies independent standard checks only to implemented QDC/QDPX-codebook, WebVTT, DOCX, and XLSX claims.
  • Both QDC and the QDPX codebook projection validate against the vendored official REFI-QDA Codebook 1.0 schema.
  • WebVTT and OOXML fixtures pass independent parsers that do not call the product import parsers.
  • Malformed QDC, WebVTT, and DOCX failure injections fail closed.
  • Complete-project QDPX, saved queries, report builder, and all partial/planned capabilities remain explicitly unproved.

Failure boundary: This proof covers implemented open-format boundaries only; it does not certify complete-project QDPX, third-party round trips, saved queries, report builder behavior, or a production service level.

OpenVerbatim is an open-source (Apache-2.0) qualitative data analysis platform for coding and analyzing interview transcripts. AI-suggested codes stay marked as suggestions until a human reviewer confirms or rejects them, and every decision is kept in an audit trail. The full feature set is available when self-hosted; there is no paid feature wall.