All selected work

fourth app

OpenIntelligence

An open-source private document intelligence app for iPhone, iPad, and Mac. It turns files and media into searchable libraries, answers with inspectable citations, and shows which execution route actually ran.

Version 5.5 on the App Store. Answers on your device, or on Apple Private Cloud Compute after you approve what is sent.

Apple IntelligenceFoundation ModelsConsent-Gated PCCSwiftUISwiftOn-Device RAGAgentic Reasoning LoopMetal SIMD4BNNS VectorsVision OCRSQLite FTS5

Screen recording

OpenIntelligence for iPhone, filmed in the Simulator on my Mac, sample documents

Scroll to play. Scroll back to rewind.

Screens

From the App Store listing for version 5.5. Scroll sideways.

  • OpenIntelligence iPhone screenshot 1
  • OpenIntelligence iPhone screenshot 2
  • OpenIntelligence iPhone screenshot 3
  • OpenIntelligence iPhone screenshot 4
  • OpenIntelligence iPhone screenshot 5
  • OpenIntelligence iPhone screenshot 6
  • OpenIntelligence iPhone screenshot 7
  • OpenIntelligence iPhone screenshot 8
  • OpenIntelligence iPhone screenshot 9
  • OpenIntelligence iPad screenshot 1
  • OpenIntelligence iPad screenshot 2
  • OpenIntelligence iPad screenshot 3

How it works

Import once, then ask. Every step below runs on the device except the one you explicitly hand to Apple.

On deviceYou decideApple, with consent5 of 6 steps stay on the device
  1. On device

    Import

    Extract, chunk, embed, and index every file and recording into vectors plus a SQLite FTS5 full-text index. Office, PDFs, scans, code, audio, video.

  2. On device

    Retrieve

    Dense vector similarity and BM25 keyword search run in parallel, merge through Reciprocal Rank Fusion, get re-ranked by a cross-encoder, then pass through MMR so the context is diverse.

  3. On device

    Pack evidence

    Expand hits to their parent sections and pack a finalised evidence envelope. Only now is the real context size known.

  4. You decide

    Choose the route

    On-Device never leaves the phone. Hybrid decides per query from the evidence actually found. PCC is used only after you have seen how much would be sent and why, and approved it.

  5. Apple, with consent

    Answer

    Apple Foundation Models on device, 27 tokens/sec on an A18 Pro. Private Cloud Compute on iOS and macOS 27, consent-gated, for the final synthesis only.

  6. On device

    Verify

    Verification gates check the answer is grounded in the retrieved text, abstain when it is not, and stamp a route badge read from the execution receipt.

The story

Why it exists and what building it taught me. Written for the homepage card, kept here in full.

Why: I wanted private document search that did more than produce a confident paragraph. I wanted to ask a question, inspect the evidence behind the answer, and know whether it ran on my device or used an Apple-managed route. That became OpenIntelligence: a native document workspace built around retrieval, citations, verification, and explicit execution policy.

Private Cloud Compute routing: Version 5.5 decides where final synthesis runs only after retrieval, when the evidence and real context size are known. On iOS and macOS 27 a long evidence-heavy request uses Private Cloud Compute only after you have seen how much would be sent and why, and given permission; on older systems every route resolves on-device. Retrieval and verification stay local, and every answer carries a route badge.

Local retrieval & verification: Parsing, Vision OCR, indexing, SQLite FTS5 and vector search, reranking, evidence packing, citation checks, and abstention all run on-device. The app still works in airplane mode through its on-device answer path and does not require an account or API key.

More than PDFs: The ingestion pipeline handles Office documents, spreadsheets, presentations, notes, scans, code, images, audio, and video, then keeps those different sources searchable inside the same private library.

Building in the Open: Since releasing it, I’ve opened up the project to the public. To manage the workflow, I set up a single-source-of-truth Notion roadmap and built custom agentic skills, meaning AI assistants can now log bugs and add features to the roadmap directly from the terminal as I build.

What I learned: The model call is the easy part. The real work is getting the right evidence, preserving source structure, handling ugly files, refusing unsupported answers, and making the whole route visible enough that a user can verify what happened.

How I Know If It Actually Works

An external benchmark, not a self-graded one

Benchmarks/ResearchFixtures/qasper_external_v1/ holds 83 questions over 40 papers drawn from QASPER (Dasigi et al., NAACL 2021), a published academic dataset I did not write the ground truth for. Each question is asked against a shared 40-document pool, so a correct answer has to be found among real distractors, not just returned from the one file that happens to contain it. The dataset is pinned by a SHA-256 in fixtures.lock.json and two independent rebuilds produce a byte-identical tree.

The harness lied to me twice, on purpose I found and fixed

Early versions reported nDCG@5 of 2.131 — on a metric defined between 0 and 1 — because chunk-level results were being scored against a document-level ground truth. A separate bug reported recall as exactly 0.0 on every run, because ground truth was keyed by filename while the scorer matched on chunk UUID, so the comparison always failed silently. Both are fixed in RetrievalStageMetrics.swift, and the fixes themselves are unit-tested against hand-worked values, not against the pipeline’s own output.

Where it stands right now

The per-stage metrics (recall@k, MRR, nDCG, precision) exist, are tested, and since 5.0 are wired into the eval harness. There’s still no headline accuracy number here, deliberately: I’d rather show the instrument and the bugs it caught than publish a score I can’t yet stand behind. That’s the standard I hold the retrieval pipeline itself to.

Shipping

Every commit, week by week, with the moments that mattered marked on it.

Commits
1,287
Active weeks
45
  1. First commit
  2. App Store release
  3. v2.1.1
  4. v3.7
  5. v4.6
  6. v5.5

Milestones

49 releases since January 30, 2026: 34 on iPhone and iPad, 15 on the Mac.

  1. 5.5

    One fix and lower prices this time.

  2. 5.0

    Documents were quietly losing parts of themselves, answers were built from a fraction of what was found, and the app rewrote your library on every launch.

  3. 1.0Mac

    First release on the Mac App Store.

  4. 4.0

    Dynamic Model Routing (On-Device & PCC): Automatically routes queries based on complexity and context size.

  5. 3.0

    Major reliability upgrade focused on hard technical documents, noisy OCR, and broken table extraction.

  6. 2.0

    New AI Hub toolbar with 5 document-aware transforms -- each uses your actual retrieved source chunks, not just the AI response text.

  7. 1.0

    First release on the App Store.

All 49 releases with their notes, on gunzino.me