UI demo · mock-backed replay
DocLocal
PDF question-answering with retrieval-augmented generation (RAG): the passages most relevant to each question are retrieved from the document and sent to the model with it. Every answer cites its source — click a citation to jump to the highlighted passage in the PDF and check it yourself.
The live build is Angular 21 and Nx against a Python/FastAPI backend on NVIDIA NIM cloud models — no model download or local GPU — with streamed answers, mid-stream cancellation, and session recovery. It is invite-only; request access from the sign-in page. That path sends document text and questions to NVIDIA.
The earlier fully in-browser path — Transformers.js embeddings in a Web Worker and WebLLM inference on WebGPU, with a SharedWorker so every open tab shares one model — is kept in the repository but is not part of the shipped app.