OpenIntelligence

Documentation · 23 documents

README and docs

Mirrored from the OpenIntelligence repository every day. View on GitHub

README

OpenIntelligence app icon

OpenIntelligence

Ask your own documents anything. Every answer cites the page it came from.

Download OpenIntelligence on the App Store Read the OpenIntelligence demo guide Read the OpenIntelligence architecture guide View the OpenIntelligence public roadmap


Drop in PDFs, Office files, code, images, audio, or video. OpenIntelligence reads them, builds a private searchable index, and answers questions about them with citations you can tap and check.

The app icon includes light and dark appearances, so it follows the device’s appearance setting.

It runs on your device. Ingestion, indexing, retrieval, ranking, and verification are all local. Nothing is uploaded to make search work. There is no account, no server of mine, and no third-party AI service anywhere in the path.

On iOS and macOS 27+, you can optionally allow Apple Private Cloud Compute to write the final answer, but only after you have seen how much would be sent and why. Every answer carries a badge showing where it actually ran, read from an execution receipt rather than from what was requested.

An answer with inline citations linking back to source excerpts Inspecting an answer's supporting evidence The live ingestion pipeline naming each stage as it runs The document library

A cited answer · inspecting the evidence behind it · the live pipeline · the library

Why it exists

Most document AI asks you to upload your files somewhere first. If those files are medical, legal, financial, or simply yours, that is the whole problem.

This is a bet that a genuinely good retrieval engine fits on an iPhone, and that an answer is more trustworthy when you can see what it was built from.

What it does

Reads real documentsPDFs with tables and figures, Office files, code, plain text, images, audio, and video. Pages, Numbers and Keynote are not readable and must be exported first. Vision handles OCR when a PDF’s text layer is unreliable.
Searches two ways at onceMeaning-based vector search and keyword BM25 run together, then fuse and re-rank. Hunting a part number and asking a conceptual question both work.
Shows its workA live pipeline view names each stage as it runs, and every claim links back to the excerpt behind it.
Explains its own vocabularyEvery figure it shows defines itself where it sits, in a plain register and a technical one. Tapping 38 TOPS says what the number is, and says that it is a per-chip rating rather than a measurement.
Answers offlineAirplane mode included.

Quality modes

Standard answers in one pass — best for lookups and direct questions. It is the only mode with a measured accuracy baseline: 0.410 exact-match across 83 cases drawn from 40 real research papers (QASPER), local-only.

Withdrawn 2026-08-21. This previously read “80% across 20 ground-truthed cases with zero hallucinations”. That figure came from a synthetic corpus written alongside its own questions, from a run now flagged in BenchmarkRuns/PROGRESSION.md as averaging 7 seconds per case, below the threshold at which generation almost certainly did not execute. The replacement is a harder external set and a worse number. See How OpenIntelligence Works, section 08.

Deep Think runs 4–8 sequential reasoning sessions over rotating context windows, passing compressed findings forward before synthesising. It stops early once it stops learning.

Maximum lifts the session ceiling for questions that span a whole library.

Neither Deep Think nor Maximum has a score against Benchmarks/rag_eval_v1.jsonl yet. Their reasoning chain was broken until mid-2026, so the question this architecture exists to answer — does more compute buy more correctness? — is still open. Please don’t cite a Deep Think accuracy figure; none has been measured.

How it works

Import    →  extract → chunk → embed → index (vectors + full-text)

Question  →  expand → hybrid search → fuse → re-rank → diversify
          →  expand to parent sections → pack context → answer → verify

Dense vector similarity and BM25 keyword matching run in parallel, merge through Reciprocal Rank Fusion, get re-ranked by a cross-encoder, then pass through MMR so the context is diverse rather than five phrasings of one paragraph. Answers face verification gates that check the response is genuinely grounded in the retrieved text before you see it.

Measured on a physical A18 Pro: 27 tokens/sec on-device, 86 tokens/sec on PCC, time-to-first-token 2.2–3.2s. The PCC figure was measured from an Xcode 27 build, which is what App Store builds have been since 5.2.

Routing, and what the badge means

The model picker is a policy, not a hint:

  • On-Device never uses PCC. Not for planning, not for synthesis.
  • PCC requests Private Cloud Compute, with a declared local fallback if a gate or quota blocks it. Since 5.2 the App Store build carries the PCC paths compiled in, so the fallback is taken only when an entitlement, availability, quota, network or consent gate actually blocks the route.
  • Hybrid decides per query, based on the evidence actually retrieved.

Cloud consent is requested only for a real, finalised evidence envelope — never at launch, never speculatively. Retrieval and verification stay local regardless of which model writes the answer.

There is no separately selectable “3B” or “20B” on-device model here, because the public SDK exposes no such selector. Apple’s larger on-device model is real and managed by the OS; no app can choose or observe it.

Documentation

The engineering docs are unusually detailed, and label every claim with how it was verified — source-read, build-checked, test-covered, or confirmed on a physical device. Where something is unproven, it says so.

If you are reading one thing, read this: How OpenIntelligence Works. Every stage of both pipelines in order, with the reason each one exists and what breaks without it. Written at two levels, plain English and technical, side by side.

Start here for everything else: Documentation Atlas

Architecture

Apple platform specifics

Honest limits

Contributing agents: RepoOS Command Center routes repository work through canonical evidence, safe edit boundaries, and required tests. The codemap names, for each feature, the code to read first and how it connects to the rest.

Codebase map

ModuleCore filesResponsibility
IngestionDocumentProcessor.swift, LayoutAwareExtractor.swiftContent extraction, Vision OCR fallback, structure recovery
ChunkingSemanticChunker.swift, ContentTaggingService.swiftContext-aware chunking, entity resolution, metadata
IndexingSQLiteFullTextService.swift, BNNSVectorDatabase.swiftSQLite FTS5 and BNNS-accelerated vector storage
RetrievalHybridSearchService.swift, ContextPackingService.swiftHybrid merge, parent-chunk reconstruction, token packing
OrchestrationRAGEngine.swift, AgenticOrchestrator.swiftRe-ranking, MMR, agentic reasoning loops
Foundation ModelsLLMService.swift, FoundationModelRoutePolicy.swiftOn-device execution, PCC escalation, routing receipts
Evidence ThreadsEvidenceThread.swift, EvidenceThreadStore.swiftLocal persistence of conversations and verification state
Storage & SyncSettingsStore.swift, WorkspaceSyncService.swiftFeature gates, StoreKit 2 quotas, iCloud workspace sync
InterfaceChatScreen.swift, DocumentLibraryView.swiftChat and library surfaces
VocabularyGlossary.swift, GlossaryViews.swiftOne definition per term in two registers, attached in place via .definedTerm(_:)
ShortcutsRAGAppIntents.swift, ScreenAwarenessIntents.swiftSiri and App Intents, resolving in-process

Building

Requirements: macOS 26 or later with Xcode 26+, iOS 26.0 deployment target, Apple Silicon (M1+ / A17 Pro+) for usable Neural Engine throughput. Xcode 27 is needed for the iOS/macOS 27 paths — Core AI embeddings and native Private Cloud Compute — which compile out below that SDK.

# iCloud sets extended attributes that break codesign — clear them first
/usr/bin/xattr -cr .

# Simulator smoke build
./scripts/build_simulator_smoke.sh

# Quality-mode benchmark matrix, 20 cases per mode
python3 scripts/run_quality_matrix.py --app <path to the built OpenIntelligence.app>

# Guard against iCloud conflict copies before any signing work
./scripts/check_icloud_conflicts.sh

The benchmark denies PCC by default so runs reproduce offline, and reports Measured separately from Unmeasured rather than scoring an empty run as a failure.

This repository lives in iCloud Drive, so builds need a -derivedDataPath outside ~/Documents, and .git is redirected through .git.nosync.

Large documents

500+ page PDFs go through a streamed, batched ingestion pipeline with page-level checkpointing, so a long import survives memory pressure and app restarts instead of starting over.

Status

Shipping on the App Store for iPhone, iPad, and Mac. 5.5 is live on both platforms, released 2026-09-30 from build 483. Developed against a public roadmap synced from the same database the work is planned in.

Private Cloud Compute shipped in 5.2, live on both platforms since 2026-09-10, and execution is confirmed on a physical device. The release pipeline enforces the claim rather than trusting it: every archive is checked with nm -u for PrivateCloudCompute symbols, against a live SystemLanguageModel count as a control so a dead binary cannot pass the gate vacuously. The version stamped on the archive decides which way the gate points: below 5.2 any PCC symbol fails the build, from 5.2 zero PCC symbols fails it, so the binary cannot silently contradict this page in either direction. Edge cases, quota exhaustion, mid-stream network transitions and background consent, are still unverified and tracked as open items rather than quietly assumed.

Releases are produced by Xcode Cloud, never from the maintainer’s Mac. This was not a preference. That Mac ran a beta macOS, every local archive stamped a prerelease BuildMachineOSBuild into Info.plist, and App Store ingestion rejected that with ITMS-90111 even though altool --validate-app returned VERIFY SUCCEEDED and processing reached VALID. Validation is not ingestion. Xcode Cloud builds on Apple’s released images, so the stamp comes out clean. The Mac now runs the macOS 27.0 release, build 26A428.

That is also where the toolchain question above is settled. Xcode Cloud runs one workflow, Default, pinned to the Xcode 27 release, build 27A266a, on macOS Latest Release, with two archive actions. Xcode 27 ships Swift 6.4, which compiles the #if compiler(>=6.4) sites in, and ci_scripts/ci_post_clone.sh fails in seconds if a 5.2-or-later version ever meets a Swift earlier than 6.4 runner.

GitHub Actions ran releases for one stretch in August 2026, as a fallback after the free Xcode Cloud allowance ran out, and was retired on 2026-08-28 once paid Xcode Cloud capacity was in place. Its workflow no longer exists.

License

MIT. See LICENSE.

Documentation

22 documents synced from the OpenIntelligence repository, 6 kept for history.

Apple Intelligence & Foundation Models Transition Plan (WWDC26 Master Blueprint)Status 2026-09-24: historical phase record, not current state. The phase-based plan this document tracked (Phase 1A–1D) has finished (AGENTS.md rule 15). For cuAppleIntelligenceTransitionPlan.mdBilling And LimitsThis document describes the billing tiers, StoreKit 2 product identifiers, and resource quota boundaries as audited in the OpenIntelligence v4.4 codebase.BILLING_AND_LIMITS.md · verified at v4.4, quota matrix re-verified at v5.3Canonical OpenIntelligence Source of Truth- The AppIcon asset catalog retains the existing light mark and includes a universal iOS dark luminosity rendition. The compiled asset catalog contains a UIAppeCANONICAL_OPENINTELLIGENCE_SOURCE_OF_TRUTH.mdDocumentation Consistency AuditDOCUMENTATION_CONSISTENCY_AUDIT.mdHow OpenIntelligence WorksCorrected 2026-09-29 against 8be003a: the GPU Acceleration profile gates Metal vector search (Efficiency and Balanced, the default, keep it on the CPU); a citatHOW_IT_WORKS.mdIngestion PipelineRead this file by section, never whole (about 65 KB): grep -n '^## ' Docs/INGESTIONPIPELINE.md, then the section a task touches.INGESTION_PIPELINE.md · source-verified at v4.6; shipped versions in Docs/SHIPPED_VERSION.jsonLimitationsOpenIntelligence ships on the App Store for iPhone, iPad, and Mac, with paid subscription tiers. It is a real product, and the limits below are the honest boundLIMITATIONS.mdOpenIntelligence Architecture AtlasRead this file by section, never whole (about 55 KB): grep -n '^## ' Docs/OPENINTELLIGENCEARCHITECTUREATLAS.md, then the section a task touches. Routes in Docs/OPENINTELLIGENCE_ARCHITECTURE_ATLAS.mdOpenIntelligence RAG Pipeline EvaluationsThis file is the single entry point for "how good is this engine, and how do I know". Nothing about measurement should live anywhere else without being linked fEVALS.mdOpenIntelligence Release NotesThis document provides a comprehensive, version-by-version breakdown of major architectural and feature updates to the OpenIntelligence Apple Intelligence-nativRELEASE_NOTES.mdOpenIntelligence roadmapThe roadmap lives in Notion, not in this file. The public view is the OpenIntelligence public roadmap, and fascinaiting.me draws its board from the same databasROADMAP.mdOpenIntelligence Study GuideThis is the course. Its job is that you can explain every component of your own app and the reason it exists, to a non-technical person, to an engineer, and at STUDY_GUIDE.mdOpenIntelligence User-Facing ChangelogThis document provides a chronological history of user-facing changes, highlighting how OpenIntelligence continuously improves transparency, speed, and reliabilUSER_CHANGELOG.mdPCC Dynamic Routing and Multi-Model Architecture Audit Specification- Phase 0: Workspace and SDK verification - Phase 1: Repository inventory - Phase 2: Runtime execution tracing - Phase 3: Apple SDK capability verificatPCC_Dynamic_Routing_Audit_Spec.mdPrivacy and Model RoutingOpenIntelligence is local-first. Extraction, OCR, embeddings, vector and lexical retrieval, evidence scoring, route planning, transcript handling, and response PRIVACY_AND_ROUTING.md · source-verified at v4.6; shipped version in Docs/SHIPPED_VERSION.jsonRetrieval Pipeline1. Import: Files enter through Apple platform document workflows. 2. Extraction: Text, layout, and metadata are extracted. 3. Chunking: Chunks are generated witRETRIEVAL_PIPELINE.md · narrative source-verified at v4.6; claims re-checked 2026-09-01
Historical documents (6)Superseded or archived upstream, kept because the repository keeps them.