Data type

Image import and OCR

Import images, run OCR, and retain region anchors.

Implemented task

Public task

Open source actions and choose Add source.

Expected outcome: A validated source enters the durable revision and processing pipeline.

ApplicationRequires Sign-in, Existing project

Declared product contract

Inputs and task boundary

Inputs

  • Supported research image

Input constraints

  • The self-host image includes English OCR data by default.
  • Other OCR languages require matching Tesseract data and configuration.

User actions

  1. Open Add source and choose an image.
  2. Run the OCR processing path.
  3. Review extracted text and region anchors in the workspace.

Outputs

  • A durable image source with OCR text and normalized page-region anchors.

Decision boundary

Where it fits

  • Images whose visible text can be reviewed after OCR.

Outside the boundary

Where it does not fit

  • Unconfigured OCR languages or an assumption that OCR needs no human review.

Current scope

Limitations

  • The self-host image includes English OCR data; other languages require additional Tesseract data and OV_TESSERACT_LANG.
  • OCR review and correction is not yet a dedicated mapping screen.

Verification

Product proof

journey

proof.m1-material-import

Open artifact
Verified
Review due
Expires
  • Eight public fixture formats enter the real source and revision pipeline.
  • Exact locators and manual coding persist across refresh.
  • Withdrawal previews impact and erases live database and stored objects.

Failure boundary: The proof does not claim M3 typed cases and variables or frame-level video analysis.

OpenVerbatim is an open-source (Apache-2.0) qualitative data analysis platform for coding and analyzing interview transcripts. AI-suggested codes stay marked as suggestions until a human reviewer confirms or rejects them, and every decision is kept in an audit trail. The full feature set is available when self-hosted; there is no paid feature wall.