Data type
Image import and OCR
Import images, run OCR, and retain region anchors.
Implemented task
Public task
Open source actions and choose Add source.
Expected outcome: A validated source enters the durable revision and processing pipeline.
Declared product contract
Inputs and task boundary
Inputs
- Supported research image
Input constraints
- The self-host image includes English OCR data by default.
- Other OCR languages require matching Tesseract data and configuration.
User actions
- Open Add source and choose an image.
- Run the OCR processing path.
- Review extracted text and region anchors in the workspace.
Outputs
- A durable image source with OCR text and normalized page-region anchors.
Decision boundary
Where it fits
- Images whose visible text can be reviewed after OCR.
Outside the boundary
Where it does not fit
- Unconfigured OCR languages or an assumption that OCR needs no human review.
Current scope
Limitations
- The self-host image includes English OCR data; other languages require additional Tesseract data and OV_TESSERACT_LANG.
- OCR review and correction is not yet a dedicated mapping screen.
Verification
Product proof
journey
proof.m1-material-import
- Verified
- Review due
- Expires
- Eight public fixture formats enter the real source and revision pipeline.
- Exact locators and manual coding persist across refresh.
- Withdrawal previews impact and erases live database and stored objects.
Failure boundary: The proof does not claim M3 typed cases and variables or frame-level video analysis.
OpenVerbatim is an open-source (Apache-2.0) qualitative data analysis platform for coding and analyzing interview transcripts. AI-suggested codes stay marked as suggestions until a human reviewer confirms or rejects them, and every decision is kept in an audit trail. The full feature set is available when self-hosted; there is no paid feature wall.