Ingestion Pipeline — source-verified at v4.6, shipped tree is v5.0
Documentation status: Source-verified on 2026-07-15 against v4.6. Not re-verified since. iOS 5.1 and macOS 5.1 are the shipped versions (both approved 2026-09-02, build 433);
Docs/SHIPPED_VERSION.jsonis the per-platform record. Corrected 2026-09-01, having said 4.9 since July; 5.1 recorded 2026-09-02. PCC Dynamic Routing does not change ingestion; indexed content remains local until a later query explicitly selects and consents to a minimized PCC synthesis envelope. Known drift as of 2026-08-05 — inCHANGELOG.mdunder 4.9 but not yet described below: all five workspace metadata writes are now atomic read-modify-writes throughcoordinatedMergeData(at:transform:), closing the race where an ingestion completing mid-sync-pass left a fully intact document on disk with no metadata row pointing at it.WorkspaceSyncServicealso no longer deletes an index for a library that still has documents. Source of truth: Codebase audit inDocs/AUDIT/, plusCHANGELOG.md4.8–4.9 for ingestion and sync. Scope: Describes shipped behavior unless explicitly labeled experimental, developer-only, or scaffolded.
This document describes the design and implementation of the import-time document ingestion pipeline, as audited at v4.6.
1. Overview
The ingestion pipeline converts raw files (PDFs, images, text documents) into searchable text segments with semantic metadata, storing them in parallel lexical and vector indexes.
flowchart TD
QLOAD[Load coordinated ingestion queue] --> QMERGE[Merge deletion-wins tombstones]
QMERGE --> QRESTORE[Read ingestion_state.json: report pages already indexed]
QRESTORE --> QDECIDE{Interrupted work remains?}
QDECIDE -- Continue --> A
QDECIDE -- Stop or Discard --> QTOMB[Persist tombstone and suppress automatic repair]
EMPTY[Metadata exists but vector index is empty] --> SINGLE[Sequential single-flight repair queue]
SINGLE --> SUPPRESSED{Library repair suppressed?}
SUPPRESSED -- Yes --> WAIT[Wait for explicit import or manual rebuild]
SUPPRESSED -- No --> A
A[Import File] --> B{File Type}
B -- PDF Size < 10MB --> C{Native Text Layer?}
B -- PDF Size >= 10MB --> STREAM[Batched Streaming Ingestion]
C -- Yes --> D[PDFKit Extraction]
C -- No --> E[Vision OCR Fallback]
D --> PRIME[Mine rough text: custom words + narrow OCR language set]
PRIME --> E
B -- Text/Markdown --> F[Direct Text Read]
B -- Image --> IMG[Structured parse, then OCR + classification]
B -- CSV --> CSVL[RFC 4180 parse to pipe rows]
B -- Office XML --> OFF[ZIP + XML: docx/xlsx/pptx]
B -- iWork --> IWORK[Always throws; see 2.6]
B -- Audio/Video --> AV[SpeechAnalyzer transcription]
D --> G[Normalizer & OCR Repair]
E --> G
F --> G
IMG --> G
CSVL --> G
OFF --> G
AV --> G
G --> H[Semantic / Structure-aware Chunking]
H --> I[Token Boundary Enforcer]
I --> J[SQLite FTS5 Storage]
I --> K[Core ML Embedding Generation]
PROFILE[GPU Execution Profile] --> E
PROFILE --> K
K --> L[BNNS Vector Storage]
J --> LOCAL[Local Search Index Boundary]
L --> LOCAL
STREAM --> STATE[Load ingestion_state.json & stable doc ID]
STATE --> S1[Process Batch of 15 Pages]
S1 --> SKIP{Already completed?}
SKIP -- Yes --> NEXT_BATCH{More Pages?}
SKIP -- No --> S2[Extract Chunks]
S2 --> S3[Vectorize Batch]
PROFILE --> S3
S3 --> S4[Store Batch to FTS5 & Vector DB]
S4 --> S5[Call db.persist & Update ingestion_state.json]
S5 --> NEXT_BATCH
NEXT_BATCH -- Yes --> S1
NEXT_BATCH -- No --> S6[Finalize Document Metadata & Await saveToDisk]
S6 --> S7[Pre-generate Suggested Questions]
S7 --> S8[Clean Checkpoints]
1.5 Import entry points (added 2026-08-27)
Everything above starts at A[Import File]. Three surfaces reach it, and all three stage files the
same way.
| Surface | Platform | Mechanism |
|---|---|---|
| Library “Add Documents” | iOS | UIDocumentPickerViewController with asCopy: true, in a sheet |
| Library “Add Documents” | macOS | MacDocumentImportPanel — NSOpenPanel.beginSheetModal(for:), no sheet |
| Finder drag-and-drop onto the library | macOS, iPadOS | .dropDestination(for: URL.self) on the library area |
All three copy into the app-managed workspace before ingesting, through
ImportedFileStaging.copyIntoWorkspace in
DocumentPicker.swift.
Importing by reference was rejected: a library pointing at files outside the workspace neither syncs
across devices nor survives the original being moved. That function is also where the
modification-date touch happens, which is what keeps WorkspaceSyncService’s 15-minute sweep from
treating a just-imported file as stale — the behaviour §5 and the canonical source of truth both
describe. A per-file failure is logged and skipped rather than aborting the batch.
[evidence_level: code_verified, confidence: exact]
macOS never uses NSOpenPanel.runModal(). runModal() starts a nested modal loop, and AppKit
refuses to start one from inside a CATransaction commit, which is exactly where SwiftUI runs
onAppear. Until 2026-08-27 all three macOS pickers did precisely that, so the panel was discarded
before it appeared: a capture that day recorded three Suppressing invocation of -[NSApplication runModalForWindow:] warnings, an _NSDetectedLayoutRecursion, and 2,064 lines of CUICatalog
window-chrome relayout from the orphaned sheet. beginSheetModal(for:) is asynchronous and has no
such restriction. [evidence_level: device_log_proven, confidence: exact, evidence_source: macOS capture 2026-08-27]
Drops are filtered before staging. Directories are excluded, because the pipeline has no concept
of a folder document and copyItem would copy one wholesale. Quota is checked before any copy, so a
drop that would exceed the document limit raises the paywall instead of half-importing.
[evidence_level: code_verified, confidence: exact]
2. Text Extraction Lanes
PDF Ingestion
- Standard Lane: Uses PDFKit to extract the native text layer if available. StructuredDocumentParser.swift is used to resolve structures like tables, lists, and headings.
- OCR Fallback Lane: If the native text layer is missing or malformed, the pipeline invokes LayoutAwareExtractor.swift to render pages as images and run local Apple Vision OCR, restoring page layout anchors.
Image Ingestion (corrected 2026-08-08)
- Standalone images (png/jpg/jpeg/heic/tiff/gif) run through
StructuredDocumentParser.parsePageImagefirst, the sameRecognizeDocumentsRequestpath PDFs use, and its output is prepended to the classification and AI description fromImageUnderstandingService. A failure falls back to the previous behaviour, so this adds structure and never removes it. - Until 2026-08-08 this path never saw the structured parser. It used
VNRecognizeTextRequest, which returns text lines in reading order and cannot express that a value belongs to a row and a column, so the same table gave cell structure inside a PDF and flat lines when photographed.[evidence_level: code_verified, confidence: exact]
Camera Capture (corrected 2026-08-08)
- Document captures run
RecognizeDocumentsRequestand prefer its structured output. Non-document captures skip it, since a photo of a pet or a landscape has no document structure to find. Failure or an empty parse leaves the flat OCR text untouched. - This path carried both defects.
CameraManager.analyzeCapturesorted its OCR observations into reading order and then joined them with a space, discarding every boundary that sort had just established, and it never reached the structured parser at all. Photographing a spec sheet produced flat text while importing the same page as a PDF produced cell structure.[evidence_level: code_verified, confidence: exact]
API currency, verified 2026-08-08. RecognizeDocumentsRequest is Apple’s current document-understanding API, introduced at WWDC25 for iOS 26; it parses tables and groups cells into rows automatically. Nothing supersedes it. VNRecognizeTextRequest is retained deliberately as the fallback for images where no document is detected, and its continued presence in these files is correct rather than debt. Sources: WWDC25 session 272, RecognizeDocumentsRequest, Recognizing tables within a document.
Text & Markdown Ingestion
- Text files, markdown notes, and source code are ingested directly. Markdown structures (headers, code blocks) are parsed to preserve hierarchical section paths.
2.5 Failure modes fixed on 2026-08-08
Four defects that produced silent degradation, meaning ingestion reported success while the content was already damaged. Recorded because each was invisible to the benchmark and each was found by reading code rather than by a failing test.
| Defect | Effect | Fix |
|---|---|---|
| OCR lines joined with a space, twice | Every ingested image became one unbroken line. Line items, dates and amounts concatenated into a single sentence; the chunker saw one paragraph and the embedding averaged the page. | Both joins preserve line breaks. |
effectiveContent chose instead of merging | A low-quality page that captured one table kept the table and discarded the recovered prose the parser had already re-OCR’d. A scanned datasheet could ship with four fifths of the page absent from chunks, FTS5 and the vector index. | Keeps structured elements and appends recovered text, deduplicated. |
usedStructuredParsing counted tables and lists only | A PDF with figures but no tables discarded every structured element, including embedded-image analysis produced at Neural Engine cost moments earlier. | Flag reflects whether structured extraction produced anything. |
| Camera captures never reached the structured parser | Photographing a document gave flat reading-order text while importing the same page as a PDF gave table cells. Same content, opposite quality. | Document captures run RecognizeDocumentsRequest and prefer its output. |
Coverage, corrected 2026-08-09. All four are now exercised by automated tests in OpenIntelligenceTests/Services/Document/Processing/. See §2.6.
The warning this section used to carry is worth keeping in its original terms, because it is the reason §2.6 exists: the evaluation corpus was 25 markdown files against 20 advertised formats, so none of the above was exercised by any automated run. A table-chunking change on 2026-08-08 destroyed 70% of retrieved table text while the benchmark reported an unchanged 16/20; it was caught only by a human noticing character counts in the running app. [evidence_level: measured, confidence: exact, evidence_source: BenchmarkRuns/20260808-retrieval-stages]
2.5.1 Two further defects, found 2026-08-09 by the fixtures in §2.6
Both were found on the first execution of the new fixture set, and neither was reachable by any test that existed before it.
| Defect | Effect | Fix |
|---|---|---|
| Word table text was extracted and then discarded | extractTextFromWordXML lifts each <w:tbl> out, leaves a bare [[TABLE_n]] marker in its place, and re-inserts the table text at that marker afterwards. But the intervening step builds its output only from <w:p>...</w:p> matches, and in OOXML a <w:tbl> is a sibling of <w:p>, never nested inside one. The marker therefore sat outside every paragraph match, never entered the result, and the re-insertion had nothing to replace. Every table in every .docx was parsed into rows and dropped, while the surrounding prose came through — so extraction looked like it worked. | Marker is wrapped in a <w:p> so it lands in the stream that step reads. Marker numbering also moved from append order to document order in the same change, but that is readability rather than a second defect: tables are replaced back-to-front, so append order labelled the last table 0 and stored its text at index 0. The pairing was reversed and self-consistent, and would have resolved correctly once substitution began working. Document order is worth having because the marker number now matches the table’s position on the page when this is debugged from a log. |
| A fully scanned PDF reported zero OCR pages | The primary structured path hardcoded usedOCR: false. It renders and recognises a page image whether or not the PDF carries a text layer, so “Vision ran” is not the signal — “there was no usable native text” is. A scan that RecognizeDocumentsRequest read cleanly reported 0 OCR pages while one that fell through to the rescue path reported 1, which is backwards from what “OCR: N pages scanned” tells the user in Document Details. | usedOCR is now isGarbled || bestAvailableText.isEmpty. |
[evidence_level: code_verified+test_verified, confidence: exact, evidence_source: DocumentProcessor.swift extractTextFromWordXML and the structured PageParseResult, IngestionFormatCoverageTests 12/12]
2.6 Format coverage, added 2026-08-09
StructuredPageContentTests pins the effectiveContent merge with no fixture and no Vision call, so that defect has a deterministic guard independent of OCR behaviour.
IngestionFormatCoverageTests drives real files through processDocument for a text-layer PDF, an image-only (scanned) PDF, a figures-only PDF, a partially legible scan, a PNG of a table, CSV, .docx and .xlsx. It asserts on extraction, not on answers:
- each table row returns with its label and its value on one line;
- a multi-line page does not collapse to one line;
- character yield holds against the characters the fixture draws, with a per-lane floor;
- the lane under test is the lane that actually ran;
- the chunk inventory has not changed shape, including that chunking did not lose what extraction already recovered.
That last group is the point. An answer score cannot see any of it: the 2026-08-08 regression held at 16/20 while destroying 70% of retrieved table text.
Fixtures are synthesised in Swift at test time rather than committed. Expectations and bytes derive from one TableSpec, so there is no hand-transcribed ground truth to drift; no binary enters an iCloud-synced repository; and the test target’s file membership does not change, so hard-boundary project.pbxproj stays untouched. .docx and .xlsx are built by a store-only ZIP writer, which works because the reader in DocumentProcessor accepts compressionMethod == 0. Reasoning and consequences in Docs/ai/DECISIONS.md.
What this does not cover, stated plainly. A page rasterised from vector text is cleaner than anything a scanner produces, so these fixtures catch structural regressions and not OCR accuracy loss on noisy input. That is an acceptable trade because all six defects above were structural, but a real scanned corpus is still worth acquiring and this set does not replace it.
iWork extraction does not work for any document current iWork produces
Not a regression; it has never worked, and the reason is structural rather than a parsing bug. extractTextFromIWorkDocument:
- inspects directories only, so a single-file
.pages— what Pages on iOS always writes — never reaches a content read at all; - for the package form, looks for members with extension
xmlortxt, while modern iWork stores content as compressed protobuf inIndex/*.iwa.
Both paths end at throw DocumentProcessingError.iWorkExtractionFailed. That is the acceptable half of the answer: it fails loudly, so nothing is indexed as a silent empty. Two tests pin both shapes so that cannot quietly change. [evidence_level: code_verified+test_verified, confidence: exact]
2026-08-21: the outward claims now match this, and the escape hatches were checked before removing
them. Five user-facing places still advertised iWork support (README.md, the App Store
description, Docs/ai/PROJECT.md, the picker caption, and the error string, which called support
“limited” when it is zero). Per the claim-audit rule the removal was evidenced in both directions,
not just by absence:
- The iOS 27 SDK on this machine exposes no iWork text-extraction API. Swept
iPhoneOS.sdk/System/Library/Frameworksfor iWork/iwasymbols and for Apple-declaredcom.apple.iwork.*type identifiers with a route to content. Nothing. - The obvious workaround is closed as well. Rendering the file with QuickLook and running the
existing Vision OCR over it reaches page 1 and no further:
QLThumbnailGeneratortakes no page index, andQLPreviewControlleris UI-only.
So implementing this genuinely means parsing Apple’s undocumented .iwa protobuf, and
ZIPArchive being file-scoped private means a new reader cannot even see the unzip helper without
an access-level change. The formats stay in the picker on purpose so the failure is reachable
and explained; removing them re-creates the defect CHANGELOG.md:162 already fixed.
[evidence_level: code_verified+sdk_verified, confidence: exact]
Spatial extraction: the word-position arithmetic, and two branches that used to fail silently
extractTextWithSpatialOrdering is the column-aware reader for every PDF page that skips Vision.
It asks PDFKit where each word sits via page.selection(for: NSRange) and groups the results into
lines, then into columns.
Until 2026-08-21 the range it asked with drifted off the word. It was built from a
hand-maintained counter that under-counted two independent ways, and the two compounded:
split(whereSeparator:) omits empty subsequences, so every run of whitespace collapsed into a
single gap while the cursor advanced by exactly + 1 per gap; and String.count counts grapheme
clusters while NSRange addresses UTF-16 code units, so ligatures drifted it further. Both
under-count in the same direction, so the range stayed in bounds and PDFKit kept returning
bounds, for different text than the word being positioned. The word appended to the line was right
and its recorded coordinates were not. The range now derives from the Substring,
NSRange(word.startIndex..<word.endIndex, in: pageString), where in: pageString is load-bearing.
[evidence_level: code_verified+test_verified, confidence: exact, evidence_source: SpatialOffsetArithmeticTests]
Two branches returned a materially worse result and said nothing, which is why a two-column extraction defect had to be diagnosed by inference across three sessions rather than read off a trace:
guard spatialLines.count > 3returns nil, sending the caller to rawpage.string, which interleaves columns by construction. Both call sites read... ?? text, so a nil return was indistinguishable from a success.- When
detectColumnBoundariesreturns empty, the page is read as a single column, sorted by Y alone, and returned non-nil. Genuinely single-column pages and multi-column pages whose boundaries were missed both land here and cannot be told apart from inside the function; in the second case a Y-only sort interleaves the columns, because a left line and a right line at the same height sort adjacent.
Both now log. The selected PageProcessingStrategy is also logged per page, and equal-Y ties break
on X ascending, because sorted(by:) is not stable in Swift and the same page could previously
extract differently on two runs. The instrumentation was the point of that change, and it is
what made the attribution below possible.
[evidence_level: code_verified, confidence: exact]
Attribution closed 2026-08-24: the column signal was destroyed before column detection ran
The damage was upstream of every check meant to catch it, which is why branch 2 above looked like the defect and was only a symptom of it.
Words were grouped into a SpatialLine on a vertical test alone. page.string on a two-column
PDF emits words row by row across the gutter, so a left-column word and a right-column word at the
same height arrive adjacent at the same Y and were merged into one line. That line’s xPosition is
the mean of its words’ midpoints, so it lands near the page centre.
detectColumnBoundaries then looks for a gap wider than 10% of the page across those means. Once
every line sits at the centre there is no gap, so it returns empty, branch 2 sorts by Y, and the
columns interleave. Neither a smarter detectColumnBoundaries nor the tie-break could have fixed
this: both operate on X positions that have already been averaged across the gutter.
DocumentProcessor.spatialLineBreak(wordBounds:currentLineBounds:currentLineY:) now returns
.none, .verticalGap or .horizontalGap, and a horizontal discontinuity starts a new line: a
move back to or past the line’s own left edge, or a forward jump wider than 1.5 × glyph height
(floored at 8pt), which is a gutter rather than a word space. Scaling by glyph height keeps the
threshold correct at any page size. Lines then stay inside one column, xPosition becomes genuinely
bimodal, and the existing column grouping and per-column top-to-bottom assembly work as designed
with no further change.
It is nonisolated static and takes only geometry, so this is the first part of the two-column path
that can be exercised without ingesting a real journal PDF on a device. Seven cases in
DocumentProcessorTests cover the gutter crossing, the left-margin wrap, ordinary word spacing,
the next row, both no-line-started guards, and that the threshold scales with type size.
A page whose words arrive across a gutter now says so, so a genuinely single-column page and a two-column page are no longer indistinguishable in a trace.
Device evidence this closes: a stored chunk read “…in locomotors are more diverse, consisting of Gi-coupled 5-HT, and 5-HTS. tion, reinforcement learning…” — lines emitted out of order, with “locomotion” severed across the join.
Withdrawn: the non-English detections are not a scrambling signal. This document was read as
non-English on 118 of 552 chunks, and that was proposed here as a free interleaving detector. The
corrected capture disproves it: with the ordering fixed and every page matching Vision’s transcript,
the distribution barely moved (126 non-English chunks to 123). Inspecting them shows why — they are
reference-list entries such as “Weissbourd B, Ren J, DeLoach KE, Guenthner CJ, Miyamichi K”, which
are surnames and initials with no function words, and any language recogniser will call those
Indonesian or Dutch. The signal was reading bibliographies, not damage.
[evidence_level: device_log_proven, confidence: exact, evidence_source: PostPostFixAgain.txt, lang distribution before and after the fix]
[evidence_level: code_verified+test_verified, confidence: exact, evidence_source: DocumentProcessor.spatialLineBreak; DocumentProcessorTests, 7 cases; device capture 2026-08-24, not yet device-verified]
2026-08-24, same evening: the real defect was line ordering in the Vision path, not columns
Two paired probes settled this in one device run and corrected two prior diagnoses.
DocumentObservation.text.transcript for page 1 reads “…dynamics that regulate diverse aspects of
motivation-related behavior. Dopamine and serotonin transiently modulate moment-to-moment behavior at
timescales ranging from sub-second to minutes…” — correct and continuous. Vision’s own reading
order was never the problem.
The text that actually won for that page — winner=layoutText on all 8 pages — emitted those same
lines in the order 2, 1, 4, 3. Adjacent pairs swapped. Pages 3 and 5 take the PDFKit spatial path,
carry its speci c ligature drops, and read correctly, which localises the defect to
LayoutAwareExtractor.
groupBlocksByLine matched a block to a line with
groups.keys.first { abs($0 - block.topY) < lineTolerance }. Dictionary.keys iterates in hash
order, so a block within tolerance of more than one line joined an arbitrary one. The grouping
therefore depended on Vision’s block order and on the per-process hash seed, so the same page could
group differently between runs. Blocks are now consumed top-to-bottom and matched to the nearest
line from sorted keys.
Nothing was read across the gutter and no text was lost. Every word was present and correctly recognised, in the wrong order — which is exactly why it passed a printable-ratio and entropy quality gate, and why the repeated non-English detections were a symptom of scrambled word sequences rather than of missing text.
Correction, same evening. The paragraph above attributed this to groupBlocksByLine’s
Dictionary.keys.first lookup. That is a real defect and it was not the cause: the capture after
fixing it was byte-identical, layoutTextChars=5717 before and after, with the same transposition.
The cause is one level up, in buildReadingOrderText:
if abs(lhs.topY - rhs.topY) > 0.02 { return lhs.topY > rhs.topY }
return lhs.minX < rhs.minX
The threshold is larger than the quantity it compares. A journal page carries roughly 55 lines,
so consecutive lines sit about 0.015 apart in normalised coordinates — under 0.02. The vertical test
therefore never fired between neighbouring lines, and minX decided their order: body text sorted by
left edge, so an indented or hyphenated line jumps ahead of the line above it.
And it is not a strict weak ordering. Lines at 0.500, 0.485 and 0.470 make both inner pairs
compare equal while the outer pair compares ordered. sorted(by:) is documented as producing an
unspecified result when given such a predicate, which is why the output was scrambled rather than
merely imprecise.
readingOrderPrecedes(lhsTopY:lhsMinX:rhsTopY:rhsMinX:) replaces it: strictly by position down the
page, left edge only for an exact tie. No epsilon is required, because groupBlocksByLine has
already collapsed a physical line into one element using a tolerance derived from real glyph height.
Applied to all three call sites in the file rather than the one that was proven, and the same
non-transitive shape in detectAndSeparateTables is now row-bucket quantisation, which keeps its
legitimate row-then-left-to-right intent while being transitive. Three unit tests pin the ordering,
the transitivity, and that the left edge never overrides a real vertical difference.
Device-verified 2026-08-24. After the comparator fix, all eight pages’ stored text matches
Vision’s transcript order. Page 1 now reads “dynamics that regulate diverse aspects of
motivation-related / behavior. Dopamine and serotonin transiently modulate / moment-to-moment
behavior at timescales ranging from / sub-second to minutes…” against the previous 2, 1, 4, 3.
layoutTextChars is unchanged at 5,717, confirming the text was only ever misordered and never lost.
The paired probes were removed once they had done their job; the three comparator unit tests replace
them, and unlike a log line they fail against the old predicate.
Correcting the entry above it: the 952d85f note described this as gutter interleaving. It is not.
That commit stands on its own — extractTextWithSpatialOrdering genuinely did merge across a gutter
and now has 7 tests where it had none — but that path handles 2 of 8 pages here and found nothing to
correct on either, and the language-detection distribution was byte-identical before and after it.
[evidence_level: device_log_proven+build_verified+test_verified, confidence: exact, evidence_source: PostFix.txt TRANSCRIPT PROBE vs PAGE TEXT PROBE across all 8 pages]
Audio and video: the failure mode is verified, transcription itself is not
say fails from an agent shell (-241), so no speech fixture can be authored there. A silent WAV was used to check the failure mode instead, and that part is settled: ingesting audio with nothing to transcribe throws in about 76 ms rather than yielding an empty document that gets indexed as a success. [evidence_level: test_verified, confidence: exact, evidence_source: IngestionFormatCoverageTests testSilentAudio, 0.076s]
One cold-start caveat, recorded because it cost a run. On the first attempt of the session SpeechAnalyzerService logged Starting analysis: silence.wav (1s) and had not returned when that run was terminated for unrelated reasons; the elapsed time was never measured, so “it hung” is an inference and not an observation. Every subsequent attempt threw immediately. The plausible reading is one-time framework or model preparation. The test caps the wait at 60 s and reports a timeout as a distinct outcome so a genuine stall would be visible rather than taking the suite down.
Still uncovered: whether transcription of real speech works at all. No automated test exercises .mp3, .wav, .mp4, .mov or .m4a with actual speech in it. That needs a committed sample or a human with an audio session.
3. Chunking & Token Gating
Semantic Chunking
- Raw text is chunked using SemanticChunker.swift. It runs adaptive windows (default size $\le 310$ words) with character overlap.
Structure-Aware Chunking
- When structured tables or lists survive the parsing phase, they are preserved as atomic chunks to prevent layout breakage, ensuring that data cells are not separated from their column headers during retrieval.
Token Limit Enforcement
- Before indexing, chunks are checked against local tokenizers (e.g.
BertTokenizer) to guarantee they are within the embedding model’s limit ($\le 510$ tokens).
3.5 Stage conservation, added 2026-08-28
Three guards now sit on this path, and they answer different questions. Confusing them is how the third one came to exist.
| Guard | Asks | Blind to |
|---|---|---|
verifyTokenizerCounts | does the token counter vary with its input? | anything after tokenisation |
verifyContentCoverage | did the finished chunks keep the extracted text? | which stage lost it |
IngestionStageLedger | did each transition conserve its input? | anything outside DocumentProcessor |
The metric that was computed and never compared
verifyContentCoverage measured two things and asserted on one. coverage is a set intersection of
unique words and had a < 90 threshold. charRatio — the same comparison by character volume — was
computed, formatted into the healthy-path debug line, and never compared against anything.
Those two metrics fail differently, and the difference is the whole point. Unique vocabulary
saturates long before content does: truncate the back half of a real document and nearly every
distinct three-letter-or-longer word still appears in the front half, so coverage stays above 90
and the function stays silent while half the document is gone. charRatio is the number that moves
in that case, and it was the number with no threshold.
It now has two bounds. Below 90% means text was dropped between extraction and the finished chunks. Above 200% means chunks are duplicating content beyond what overlap explains; configured overlap puts a healthy document around 115–125%, so the ceiling has wide margin and still catches the duplicate-import shape. Both numbers are reported on every warning whichever bound trips, because high word coverage beside low volume is the truncation signature specifically, and is a different fault from both being low.
[evidence_level: code_verified, confidence: exact, evidence_source: DocumentProcessor.verifyContentCoverage before 2026-08-28 — charRatio assigned once, interpolated into one Log.debug, no comparison anywhere in the file]
Localising a loss to its stage
verifyContentCoverage is a single end-to-end comparison across four stages: the structure-aware
path or the semantic fallback, then sanitizeProcessedChunkMetadata, then
enforceTokenLimitOnChunks. It can say text was lost and cannot say where, which leaves a bisect by
hand.
IngestionStageLedger.swift records each transition separately and gives each one a band, so a loss names its own stage:
| Transition | Band | Why |
|---|---|---|
| extraction → chunked | ≥ 0.98 | overlap repeats text so the ratio may exceed 1.0; normalisation trims a little |
| chunked → sanitized | exactly 1.0 of characters | that pass rebuilds ChunkMetadata and never reads or writes chunk.text |
| sanitized → token-limited | at least 0.99 of words | that pass splits rather than truncates, but it is not character-preserving |
The last row is the sharp one, and it is checked in the wrong unit if you are not careful. A split raises the chunk count while conserving content, so a count alone cannot tell a healthy split from a truncation — both move the number and only one keeps the text.
But splitOversizedChunkByTokens separates on .!?\n using components(separatedBy:), which
discards every separator, then rejoins with ". ", injects [Part N] markers, and drops
blank-line runs entirely. No word is lost — both flush paths are present — but the character count
moves in both directions. This band was written as exact-characters on 2026-08-28 and corrected the
same day, before it ever ran on a real document: it would have fired on the first document
containing an oversized chunk, which is a false alarm inside the instrument built to stop false
confidence. Words are what the stage conserves, so words are what is checked.
Noted and not changed: that the splitter rewrites ! and ? to . and collapses blank lines is
a real property of the text that reaches embedding. Changing it changes chunk content and therefore
every vector, so it is a separate decision with its own blast radius, not a drive-by fix.
The ledger also flags a run where every chunk came out the same length past three chunks. A stage
that has stopped varying with its input can still conserve characters in aggregate, so no ratio sees
it; this is the same argument verifyTokenizerCounts makes one layer down, applied to the chunker.
Every reading counts chunk.text directly and never chunk.metadata.characterCount. A defect in a
counter is one of the things being hunted, and auditing a stage with the count that stage recorded
makes the audit agree with the bug.
Not covered: chunking → embedding and embedding → index. Both live in RAGService.swift, outside
the Services/Document/** edit boundary the RepoOS router sets for ingestion work. They are the
natural next pass.
[evidence_level: code_verified + test_verified, confidence: exact, evidence_source: IngestionStageLedgerTests, 10 cases]
4. Dual Index Storage
Once chunks are generated and validated, they are written to two separate storage engines:
- Lexical Index: Stored in SQLite FTS5 via SQLiteFullTextService.swift. BM25 column weights prioritize section titles and entity tags.
- Vector Index: Dense query vectors (384-dimensions) are generated using a local Core ML model (
EmbeddingModel.mlpackage) and stored in BNNSVectorDatabase.swift using Cosine Similarity. Both indexes are isolated bycontainer_idto enforce library boundaries.
5. Performance Optimizations & Checkpointing
Zero-Copy CGImage Processing
- To reduce memory allocations and CPU overhead during structure-aware parsing and OCR fallbacks:
- Bypasses raw pixel drawing and PNG serialization passes.
- Converts preprocessed
CIImageinstances directly to rawCGImagepointers utilizingCIContext. - Vision’s
RecognizeDocumentsRequestandRecognizeTextRequestperform analysis directly on the rawCGImagememory block, accelerating extraction by 30%+ and avoiding OOM memory spikes.
macOS page rendering, corrected 2026-08-29
The zero-copy claim above was true of the iOS branch and false of the macOS one. renderPDFPageAsImage
has always had two implementations, and only the UIKit half did what this section describes.
What the macOS branch used to do. NSImage(size:) + lockFocus(), then tiffRepresentation ->
NSBitmapImageRep(data:) -> .cgImage. Two costs, both macOS-only and both invisible in the logs:
lockFocusis deprecated, andAppKit/NSImage.hgives the reason: it “is incompatible with resolution-independent drawing”. It sizes its backing store from the display’s scale factor.tiffRepresentationthen encoded that whole raster to uncompressed TIFF in memory andNSBitmapImageRep(data:)decoded it straight back. A full CPU serialize/deserialize round-trip, per page. Exactly the “PNG serialization pass” the section above says the pipeline avoids.
Measured, not estimated. A standalone AppKit probe on this host (NSScreen.main.backingScaleFactor
= 2.0) reproducing the old code path at the default scale = 5.0:
| requested | actual raster | per-page bytes | |
|---|---|---|---|
NSImage.lockFocus + TIFF | 3060x3960 | 6120x7920 (4.0x the pixels, 16 bits per component) | 370 MB, encoded then decoded |
CGBitmapContext (now) | 3060x3960 | 3060x3960 (1.0x) | 46 MB, no serialization |
Eight times less raster, and roughly 740 MB of per-page memory traffic removed. The wall-clock effect on a real document has not been measured yet; do that on device before quoting a speedup.
What it does now. A CGBitmapContext at explicit pixel dimensions, opaque, drawn into directly
and read out with makeImage(). No oversize, no serialization, and no mutation of
NSGraphicsContext.current, which lockFocus touches and which is process-global state shared with
every concurrent Vision operation. A CGBitmapContext has the same bottom-left origin as the
lockFocus context it replaces, so the existing transform and page orientation are unchanged.
[evidence_level: measured, confidence: exact, evidence_source: standalone AppKit probe on this host 2026-08-29; DocumentProcessor.swift renderPDFPageAsImage; MacOSX27.0.sdk AppKit/NSImage.h:294; macOS Debug build succeeded 2026-08-29]
Per-page render timing is now logged. The line reads
Rendered PDF page at WxHpx (N DPI) in X.Xms via <backend>, CPU raster[, post-process attached].
Two things about it matter:
- The old line ended in
[GPU-accelerated], which was a string literal selected by theuseGPUForPDFRenderingflag, not a measurement. The raster is CPU work on both platforms, and the Core Image chain that flag selects is lazy, so no GPU work has run when the line is emitted. A tester reading[ANE]and[GPU metal]in the logs while Activity Monitor showed only CPU was reading a claim about a setting, and it pointed the investigation away from the real bottleneck. - Paired with the existing
Page N: OCR extracted ... (Y.YYs)line, it splits render cost from OCR cost per page, which is the measurement that did not exist when a 210-page PDF took five hours on a Mac and nothing in the log said which stage owned the time.
Known remaining gap, not fixed here. PageComplexityAnalyzer.renderPageForAnalysis is
#if canImport(UIKit) ... #else return nil #endif, and both production call sites reach it through
analyzeBatch, which passes no pageImage. So Phase 4 of the complexity analysis — the Vision
refinement that merges imagePresence, tablePresence, chartPresence and figurePresence — never
runs on macOS, and isMixedModeScanned is permanently false there. Vision merges with max(...), so
absent it macOS scores pages as less visual than iOS would, never more. Tracked separately; it is a
classification-accuracy defect rather than a performance one.
[evidence_level: code_verified, confidence: exact, evidence_source: PageComplexityAnalyzer.swift:1045-1072 and :294-307; DocumentProcessor.swift:3323 and :4049]
The live OCR request, corrected 2026-08-30
The primary OCR path uses RecognizeDocumentsRequest, not VNRecognizeTextRequest. An audit on
2026-08-29 asserted the opposite and was wrong. extractStructuredPDFContent routes every OS 26+
device — which is every shipping device — through extractWithStructuredParsing and on to
StructuredDocumentParser, which builds RecognizeDocumentsRequest directly and never calls
OCRConfiguration.configureRequest. The older VNRecognizeTextRequest survives in six places: the
text-layer validation sample, camera capture, LayoutAwareExtractor, and pre-26 fallbacks.
Two consequences, both fixed here.
The language narrowing added on 2026-08-29 never executed. It was placed in
extractTextFromPDFWithPages, the pre-26 fallback branch, and it configured OCRConfiguration,
which the live request does not use. Verified against a live trace: OCR languages narrowed
appeared zero times across every capture. Detection is now hoisted into
extractStructuredPDFContent above the version branch, made idempotent, and threaded through
parsePageImage(_:pageNumber:customWords:recognitionLanguages:...) into
request.textRecognitionOptions. automaticallyDetectLanguage is now derived —
(recognitionLanguages == nil) — rather than hardcoded true beside an explicit thirteen-language
list, which asserted both halves of a contradiction.
minimumTextHeightFraction was 0.0 on the live request too, with the comment “Detect all text
sizes”. It is now 0.004.
The value is a fraction of image height, so it is invariant to render scale: a 6 pt glyph on a
792 pt page is 0.0076 of the image whether rasterised at 144 or 432 DPI. One constant is therefore
safe across every render setting. 0.004 corresponds to roughly 3 pt text on US Letter and
2.4 pt on A5 — below the smallest legible print in any real document.
Apple’s header for the identically-named property on VNRecognizeTextRequest is the primary source
for what 0.0 costs: “the image gets processed at the highest possible resolution with no
downscaling. With that the processing time will be the longest and the memory usage the highest.”
The new API’s property carries the same name, type and documented meaning.
[evidence_level: code_verified, confidence: exact, evidence_source: StructuredDocumentParser.swift RecognizeDocumentsRequest configuration; DocumentProcessor.extractStructuredPDFContent branch; iPhoneOS27.0.sdk VNRecognizeTextRequest.h:77 for the 0.0 semantics; live pipeline_trace.log showing zero narrowing lines before this change]
Direction of caution. Text below the floor is not recognised at all and no later stage can tell it was there. This is a number to lower if it proves wrong, never to raise casually. No wall-clock figure is claimed; the per-page render timing already logged is what will measure it.
OCR language narrowing, added 2026-08-29
OCRConfiguration.recognitionLanguages carries thirteen languages — including ja-JP, ko-KR,
zh-Hans and zh-Hant — and until now every VNRecognizeTextRequest in the codebase received all
of them, alongside automaticallyDetectsLanguage = true.
The documented reason to narrow is accuracy, not speed, and an earlier draft of this section had that backwards. Apple’s header describes what the flag does:
“Language detection will try to automatically identify the script/langauge during the detection and use the appropiate model for recognition and language correction… it will for instance determine if text is latin vs chinese so you don’t have to pick the language model in the first case. But as the language correction cannot always guarantee the correct detection, it is advisable to set the languages, if you have domain knowledge of what language to expect. The default value is NO.”
Two things follow. Vision was already identifying the script and selecting a model, so “thirteen models were being loaded” was never established and should not be claimed. What Apple does state is that the auto-detection can be wrong, and a wrong detection applies language correction against the wrong lexicon — which corrupts text rather than merely slowing it down. Setting the languages is the documented remedy.
A speed benefit is likely as well: fewer candidates to correct against, and the per-request script detection no longer runs at all. Apple does not quantify it, and neither does this document.
During PDF ingestion that knowledge is already available. Phase 1.5 of
extractTextFromPDFWithPages already joins the first 50 pages of PDFKit text to mine
document-specific vocabulary; phase 1.6 now runs LanguageDetectionService.analyzeDocument over
that same string, once per document, and OCRConfiguration.narrowedRecognitionLanguages(matching:)
turns the result into a subset of the curated list. configureRequest takes it as an optional
languages: parameter and disables automaticallyDetectsLanguage when it is present.
The function is deliberately reluctant. It returns nil — meaning keep all thirteen — whenever
narrowing is not clearly safe: a primary detection at or below 0.8 confidence, an undetermined
language, an empty rough text layer (which is what a garbled-text-layer document produces), or a
detected language the curated list does not cover at all. The asymmetry is the whole design.
Narrowing wrongly costs recognition accuracy on text that is never recovered; keeping the full list
only costs time. A future edit that makes this eager would look like an optimisation and be a
regression.
Chinese narrows to both scripts. Coarse detection cannot separate Hans from Hant, and choosing wrong there loses the text outright, so the prefix match keeps both. Eleven models dropped, not twelve.
Both outcomes log at info, not debug:
[DocumentProcessor] OCR languages narrowed 13 -> 2 [en-US, en-GB] (detected English @ 0.98)
[DocumentProcessor] OCR languages NOT narrowed, keeping all 13 (detected Unknown @ 0.00)
The second line matters more than the first. A silent fallback to the most expensive setting is the failure shape this repository keeps rediscovering, so the expensive path has to say so.
Call sites outside PDF ingestion — the camera capture path, IntelligentDocumentProcessor — pass no
languages: argument and are unchanged, because there is no prior text layer there to detect from.
[evidence_level: code_verified, confidence: exact, evidence_source: iPhoneOS27.0.sdk VNRecognizeTextRequest.h:72; OCRConfiguration.narrowedRecognitionLanguages; DocumentProcessor phase 1.6; OCRLanguageNarrowingTests, 12 cases]
No performance figure is claimed. The correctness argument stands on Apple’s own text; the speed argument is a plausible prior that only a device trace can settle, and the log line above exists so that trace can attribute it.
Page-Level JSON Checkpointing & State Persistence
-
Scope, clarified 2026-08-29: this applies only to
importLargePDFStreamed, which is entered whendocumentType == .pdf && fileSizeMB > 10(RAGService.swift:5666). Every PDF at or under 10 MB, and every non-PDF of any size, takes the non-streamed path and has no checkpoint at all. Batch granularity is 15 pages, so a resume can redo up to 14 pages. -
A restored item now reports the progress it kept, fixed 2026-08-29. Queue restoration used to set
resumedItem.progress = nilalongsidestage = .paused, so a document that had already indexed 150 of 210 pages came back displaying nothing. The work was never lost, but the UI gave a user every reason to believe it was — and cancelling is the one action that genuinely destroys it, because discarding a queue item deletes the checkpoint directory while a restart does not.DocumentProcessor.restoredIngestionProgress(for:knownPageCount:)reads the state file during restore and supplies both the detail line and the progress fraction.Two properties of that reader are deliberate and worth keeping. It never parses the document:
restorePersistedIngestionQueueIfNeeded()is@MainActorand runs during launch, so the page total is taken from the interrupted run’s own persistedmetrics.pageCountrather than from reopening a potentially very large PDF on the main actor. Its only I/O is onestatand one small JSON read. And it distinguishes a corrupt state file from an absent one: no checkpoint, a not-yet-started checkpoint (lastCompletedPage == -1), and a fingerprint that no longer matches all returnnilsilently because all three genuinely mean “nothing preserved”, but a state file that exists and will not decode logs a warning, because that means a resume is about to redo work it did not have to. When the page total is unknown the fraction staysniland the bar stays indeterminate rather than showing a fabricated 0%.[evidence_level: code_verified, confidence: exact, evidence_source: RAGService.swift restorePersistedIngestionQueueIfNeeded; DocumentProcessor.swift restoredIngestionProgress; RestoredIngestionProgressTests, 14 cases] -
To prevent data loss and avoid reprocessing from page 1 during large document ingestion:
-
Each page’s intermediate
PageParseResultis serialized to a Codable JSON format. -
Ingestion state and progress are tracked in a session-level
ingestion_state.jsonfile inside the checkpoints folder:localCacheDir()/IngestionCheckpoints/<docFingerprint>/ingestion_state.json -
This state file persists a stable
documentId(ensuring chunks are indexed under the same ID on resume),lastCompletedPageindex, and accumulated counts (chunks, words, chars). -
If the queue is paused or the app restarts mid-ingest:
- The engine restores the stable
documentIdand counts fromingestion_state.json. - The loop skips rendering, parsing, embedding, and storage tasks for any page batch where
endPage <= lastCompletedPage.
- The engine restores the stable
-
After successfully committing each page batch to the FTS5 index and vector DB,
db.persist()is called to flush vector changes, andingestion_state.jsonis updated atomically. -
Upon successful document indexing completion or user queue item discard, the temporary checkpoint directory (containing page checkpoints and the state JSON) is deleted.
Authoritative Queue Dismissal & Automatic Repair
- Stop/X and paused-item discard add a bounded tombstone for each removed queue ID to the coordinated
ingestion_queue.jsonstate. The field is optional while decoding, so older queue files remain readable. WorkspaceSyncServicemerges tombstones before queue items. A matching stale local or shared item is filtered out, and a tombstone-only file is retained so an empty local queue can still defeat a stale iCloud snapshot.- Empty-vector self-healing requests enter one sequential in-process scheduler. Stop/Discard persists per-library suppression in local preferences and the rebuild checks it before each safe document stage.
- A rebuild that has already removed a document completes the matching re-add before yielding, preventing cancellation from leaving catalog metadata partially deleted. A later explicit import or manual rebuild clears the suppression.
- The tombstone history is capped at 512 newest IDs; a future explicit import receives a new ID and is not blocked by an older dismissal.
[evidence_level: code_verified, confidence: high_pending_runtime_validation, evidence_source: RAGService.swift, WorkspaceSyncService.swift, IngestionQueueOverlay.swift]
Streamed Ingestion Append Support
- To support progressive SQLite indexing during batch-based streaming ingestion:
SQLiteFullTextService.swiftimplements optionalappendparameters onstore,storePages, andstoreChunksto bypass default UPSERT deletes.- Document FTS5 text is combined iteratively with existing content, pages are inserted sequentially, and all batch chunks are appended.
- Page index numbers are aligned dynamically using the batch
pageRangebounds offset to prevent page mapping collisions.
5b. Token counting, and the invariant that guards it
The token counter must be asked whether it counts. DocumentProcessor.countTokens returns
try tokenizer.encode(text:addSpecialTokens:).count. If the bundled tokenizer carries a padding
block, that call returns the pad width for every input, so the counter becomes a constant while
still looking like a measurement.
That is not hypothetical. Until 2026-08-17 both embedding_tokenizer.bundle and
reranker_tokenizer.bundle carried "padding": {"strategy": {"Fixed": 128}} alongside
"truncation": {"max_length": 128}, and three defects followed from that one block:
- 55% of all library content never reached the embedder. Both compiled models declare input
shape
[1, 512]and both Swift providers pad to 512 themselves, so the tokenizers were the only component capped at 128. Measured with the real WordPiece vocab over 139 live chunks: median chunk 273 tokens, and 125 of 139 (90%) truncated. - The
safeTokenLimitguard at 430 could never fire, becausecountTokensalways returned 128. The chunker had no working measure of its own output for the entire life of the feature. - Mean pooling averaged
[PAD]into every vector, because the provider buildsattentionMask = Array(repeating: 1, count: inputIds.count)over already-padded ids, so the mask marks padding as real content.
Deleting the padding block fixes all three, because every consumer already pads to its own fixed
width. Truncation is now 512, matching the models.
Why it survived so long, and what now guards it. The logs read
maxTokens=128/430 on 3,910 of 3,910 recorded ingestions. A constant that sits inside a
plausible range looks exactly like a working measurement, and nothing asked whether it should vary.
DocumentProcessor.verifyTokenizerCounts now asks at load: it encodes a one-word string and a
180-word string and logs an error if they measure the same. The check is behavioural rather than
configuration-based, so it holds however a future tokenizer expresses padding and does not require
parsing tokenizer.json. It logs rather than traps, because a wrong token count degrades retrieval
silently but does not corrupt data, and refusing to ingest would be the worse failure.
Existing libraries do not benefit until re-embedded. Old vectors remain valid 384-dimensional
MiniLM embeddings of a truncated, padding-diluted input. Nothing detects the change, because
KnowledgeContainer persists only embeddingProviderId and embeddingDim and neither moves. See
the Notion row “Existing libraries keep truncated vectors because nothing detects the embedding
change”.
[evidence_level: measured, confidence: high, evidence_source: real WordPiece tokenization of 139 live chunks; 3910 constant log lines; BenchmarkRuns/tokfix showing 386 to 430 after the fix]
6. Predictive Self-Tuning & Dynamic Config Optimization
To prevent wasteful document rebuild loops (re-extraction and re-embedding), the pipeline implements Predictive Self-Tuning and Non-Destructive Adjustments to optimize the RAG parameters before ingestion starts:
Predictive Document Pre-Scan
When a document is selected for import and the container’s autoAdaptDimension flag is enabled:
- Sample Extraction: The system extracts a raw text preview from the first 10 pages of the document (or the first 10,000 characters).
- Feature Analysis: The
LibraryIntelligenceCenteranalyzes the preview for structural, linguistic, and content signals:- Code/Math Content: Regex and keyword patterns scan for syntax or equations.
- Layout Structure: Measures list patterns, table structures, and multi-column density.
- Language Properties: Classifies vocabulary richness, multilingual complexity, and technical jargon density.
- Pre-Ingestion Adaptation: Based on these signals, the engine resolves the optimal
ChunkingPlan(e.g.densePrecisionstrategy,300word window for structured text/code) before processing page 1. - Dynamic Container Tuning: The container’s active
chunkingDirectiveis updated to.autowith these parameters. The ingestion pipeline processes the current document and all subsequent imports using this custom-tuned configuration immediately, eliminating the need to re-process the file.
Non-Destructive Configuration Shifting
- Chunking Strategy & Window Shifts: Since chunk size variations do not break vector math, changes to the chunking strategy, window window size, or overlap are applied instantly and silently to the container configuration. The database continues to perform cosine similarity searches over existing mixed chunks, avoiding the CPU/battery drain of full database rebuilds.
- Embedding Provider Shifts: Changes to the embedding provider or vector dimension (e.g., from 384D Core ML to 512D Contextual) change the mathematical vector space. Mixing dimensions will crash similarity search; therefore, embedding shifts are blocked during active ingestion and require a full database rebuild to guarantee consistency.