Production
Live · v1 Architecture

Local RAG on Apple Silicon

EQBook is a private, on-device AI research workspace for macOS — NotebookLM-style Q&A powered entirely by local MLX models. Documents are ingested, chunked, embedded, stored, retrieved, and answered with citations on the local machine. No cloud, no accounts, no document content leaving the device.

PlatformmacOS · Apple Silicon StatusLive · v1 InferenceLocal · MLX PatternRAG with citations
EQBook on-device research workspace interface

01Problem

A “chat with your documents” tool is genuinely useful, but the usual implementation sends document text to a hosted model and vector database. For sensitive material — research, legal, personal, or proprietary content — that is often unacceptable. The problem was to deliver the full RAG experience while guaranteeing, architecturally, that document content never leaves the user’s machine.

02Product

A workspace where you add documents and ask questions and get cited answers grounded in your own material. Every stage of the pipeline runs locally, so privacy is a property of the system rather than a promise in a policy.

03Architecture

The dotted boundary of the whole pipeline is the device. Nothing in this diagram makes a network call for model inference or storage.

04Engineering decisions

Why local inference?

The entire value proposition is that sensitive documents stay private. Local inference makes that structural: there is no server that could log, leak, or be subpoenaed for the content.

Why MLX?

MLX is built for Apple Silicon’s unified memory and acceleration, which makes running capable models on a laptop practical. The tradeoff is fitting within device RAM and latency budgets instead of unbounded cloud compute.

Why a local vector store?

Embeddings are derived from document content, so they’re as sensitive as the documents. Keeping the index on-device keeps the privacy boundary intact end to end.

Why RAG instead of a bigger model?

Retrieval grounds answers in the user’s actual sources and enables citations, which matters more for a research tool than raw model size — and it keeps a smaller local model useful.

How are answers kept grounded?

Responses are built from retrieved passages and cite their sources, so the user can trace every claim back to their own documents rather than trusting an ungrounded generation.

What are the privacy boundaries?

Documents, embeddings, the index, and inference all stay on the device. There is no account system and no cloud dependency for the core workflow.

05Implementation

  • Ingestion & chunking — documents are parsed and segmented into retrievable passages.
  • Embeddings — generated locally and stored in an on-device vector index.
  • Retrieval — top-k relevant chunks are selected for each query.
  • Generation — a local LLM run via MLX produces cited answers from retrieved context.

A reference implementation of this pattern — using Qdrant and Ollama — is available as an open-source example.

06Production

  • Packaged as a native macOS application for Apple Silicon.
  • No backend infrastructure — the full pipeline runs on the user’s machine.
  • Model assets are managed locally; there is no inference API dependency.

07Current status

Live · v1 released   Shipped for macOS on Apple Silicon. The underlying local-RAG pattern is also documented and demonstrated in open source.