← Nightjar Browser

Nightjar Browser — Known Limitations

What the product cannot do today, why, and what each would take to lift.

This is about capability limits. For the code-audit backlog (unverified download digests, the org-name question, the UI-thread scrape) see FIXES.md. Anything listed here is a deliberate current boundary, not a bug to be filed.

Last reviewed: 2026-08-25.


1. No OCR — scanned PDFs yield nothing

pypdf reads a PDF's existing text layer. A scanned or photographed document has no text layer, so extraction returns empty and the attachment is refused with:

No selectable text found. This PDF is likely a scan, and there is no OCR available offline.

A PDF exported from Word, a browser, or most report tools has a text layer and works fine. A PDF that is a picture of a page does not.

Feasibility (probed 2026-07-28, works end to end): every piece needed is already on a Windows 11 machine, with no new download:

Piece Status
pypdfium2 — renders any PDF page to a bitmap already a dependency
Windows.Media.Ocr — offline OS OCR engine ships with Windows 11; en-US present

A synthetic scan (text PDF flattened to an image-only PDF) round-tripped through render → OCR and came back verbatim, including the figure 8,450.

What it would take: an engine/ocr.py that renders empty pages with pypdfium2 and calls the OS engine, plus a fallback hook in extract_attachment. The one open decision is how to reach WinRT from Python: the winsdk pip package (in-process, ~0.2–0.4s/page, adds a dependency) or a PowerShell shell-out (no new dependency, ~1s/page of process startup). A per-page fallback rather than all-or-nothing would also catch documents that mix real text pages with scanned inserts.

Deferred by decision, not blocked.


2. No image input, even on multimodal models

Attaching an image is refused. The refusal names the loaded model, because the reason differs:

The second case is the app's limit, in three places:

  1. inference/manager.py builds the llama-server command with no --mmproj, so no vision projector is loaded.
  2. query_completion sends "content": prompt as a plain string. Images require the OpenAI content-parts array with image_url data URIs.
  3. inference/registry.py entries download only the main GGUF; no mmproj companion file is fetched for any model.

grep for mmproj|image_url|libmtmd across the codebase returns nothing.

Do not tell users "switch to Gemma and it will work" — it will not, until the above is wired. engine/attachments.py:image_rejection_reason() already takes a vision_enabled flag: flip it to True once the projector is loaded and the messages correct themselves.


3. Conversation memory is bounded by the context window

(Resolved 2026-07-28 — previously there was no memory at all.)

Prior turns are now sent with each question, so follow-ups and references resolve. What remains is a limit rather than a gap: history competes with retrieved documents for one window, and engine/context_budget.py divides it — roughly 70% to retrieval, the rest to history, with a floor under retrieval so an attached document is never squeezed to nothing.

Consequences worth knowing:

The window is no longer fixed at 8192; it is sized per model and machine (see §11), so on capable hardware there is far more room than there was.


4. Replies stream, but not one token at a time

(Resolved 2026-07-31 — previously the whole reply appeared in one go after a silent wait of many seconds.)

query_completion takes an on_delta callback; given one it sets "stream": true and reads the SSE body, so text appears as it is generated. Measured on Qwen 1.5B on CPU: first text at 0.28 s against 4.83 s for the complete reply, and 100 update frames across a 12.4 s answer.

What remains is a deliberate limit rather than a gap. Frames are coalesced on a ~100 ms timer (MainWindow.STREAM_FLUSH_SECONDS) instead of being forwarded per token, because delivery goes through evaluate_javascript, which is synchronous and pumps the Tk event loop while it waits — roughly 10 ms of blocked UI thread per message. At 47–162 tokens/second a frame per token would spend more time painting than generating and the browser would stop responding.

Consequences worth knowing:


5. Retrieval is keyword matching, not semantic

engine/rag.py is a hand-written TF-IDF + cosine-similarity index. It matches on shared words, with no notion of meaning.

Mitigated but not solved: the opening chunk of every source is always included (anchor_indices()), so an attachment is never dropped entirely just for missing the query wording.

What it would take: a local embedding model. That means a second model download and a second server process, or an ONNX embedder via the already installed onnxruntime.


6. Attachment limits

Limit Value Where
Max single file 32 MB MAX_FILE_BYTES
Max extracted text 400,000 chars, then truncated MAX_TEXT_CHARS
Retrieved context ~70% of the free context budget engine/context_budget.py

The model never sees a whole document — only the retrieved chunks. "Summarize this 200-page PDF" summarizes the parts retrieval selected, not the PDF. For broad questions over long documents the answer will be partial, and it will not say so.

The budget is now measured in tokens against the model's real window rather than a fixed chunk count. The previous fixed count could overflow outright: 18 chunks is roughly 7,500 tokens against a 7,168-token budget at 8k context.

Unsupported formats: DOCX, PPTX are refused. They are ZIP containers of XML, not text, and need a parser (python-docx) that is not a dependency. Legacy .doc/.xls/.ppt are worse and effectively out of scope; .xls says to save the file as .xlsx or CSV.

Supported: 49 extensions — PDF (with a text layer), HTML, TXT, MD, CSV/TSV, Excel (.xlsx, .xlsm, every sheet), JSON, XML, YAML, and 27 source-code extensions.

Calculations on spreadsheets. For a data question about an attached CSV, TSV or Excel file (a total, average, count, highest or lowest, difference or lookup), the model only writes a query and Nightjar computes the answer over every row, showing its working. Grouped breakdowns ("total per region") are not supported yet. When a question cannot be calculated, the answer is the model reading the table as text, and a note before it says so and why. Nightjar refuses to calculate from a file cut at the size limit, or from a workbook whose formulas were never calculated by Excel, since either would give a wrong total.


7. Attachments persist until removed

A file stays attached for the whole session, across every subsequent question, until removed from the tray, cleared, or dropped by "New Chat". This matches how chat attachments behave elsewhere, but it means a file attached for one question keeps consuming the retrieval budget for later, unrelated ones.


8. No Linux build, and macOS is newer than Windows

Linux: none, and no near-term path to one. The browser is WebView2 via .NET interop (pythonnet, tkwebview2) on Windows and WKWebView on macOS. Neither has a Linux equivalent here, and nothing is planned.

macOS: Apple Silicon only, and much less exercised. There is a build -- build_mac_app.py, produced by .github/workflows/build_macos.yml -- with the inference runtime sealed into the bundle and Metal for acceleration. Intel Macs are excluded deliberately: no unified memory and no Apple Silicon GPU means CPU inference at single-digit tokens per second.

Do not read that as parity. The Windows build is the one with users on it; the macOS build has had far less real use, and features are likelier to be missing or half-wired there. A separate PySide6/QtWebEngine port that would unify both platforms is planned but not started.

The OCR path in §1 is absent everywhere, and would be Windows-specific if it existed.


9. Ad blocking yields to playback, by design

The governing rule: always show the video, even if that means showing the ad. Blocking may fail by letting an advert through. It may not fail by costing the user the content they came for.

Which engine this describes. The default blocker is Brave's adblock-rust, matching every request against EasyList and EasyPrivacy inside the browser; it applies this rule through video_safety_exceptions(), layered over the lists so the exceptions win. The steps below are the hostname blocker's, which is the fallback when adblock-rust cannot load -- the same rule, spelled out.

is_blocked() enforces it in three steps:

  1. An explicit allowlist.txt / DEFAULT_ALLOWLIST entry wins outright.
  2. A host that is first-party to an open page, or that names itself as content infrastructure (api., cdn., manifest., registry., hls., or a CDN token in the domain such as phncdn), is blocked only if our own curated list names it. A downloaded community list is not trusted here.
  3. Everything else gets the full blocklist.

Step 2 exists because registry.api.cnn.io ships in the StevenBlack hosts list as a tracker and is also the API CNN's player calls to resolve a video to its media. Blocking it did not remove an advert; it meant CNN videos never played. Note it is on cnn.io while the page is cnn.com — a sibling domain, so first-party matching alone does not save it, which is why the infrastructure heuristic is there too.

The cost, stated plainly: a tracker hiding behind an api. or cdn. label that only a community list knows about will load. This is mitigated by keeping known analytics and identity vendors in the curated list (amplitude, mixpanel, segment, hotjar, clarity, permutive, lotame and ~30 more), since curated rules still apply to infrastructure-labelled hosts. Verified: 27 ad/tracker hosts blocked, 0 leaks; 15 content-delivery hosts allowed, 0 breakage.

Four further protections come from the same rule:

What this still cannot fix: a CDN hostname that reveals nothing and is not on the explicit list (ev-h.phncdn.com was caught by the phncdn token, but something like x7f2.example.net would not be). If a site breaks, the blocked-domain list in the shield drawer names the culprit and allowlist.txt is the remedy.

Related failure mode, now mitigated: the filtering proxy carries all browsing, not just adverts, so a dead listener meant every page failed with a proxy error. The accept loop used to break on any OSError. It now retries and rebinds the same port (which cannot change — it is fixed in WebView2's launch arguments), the shell re-checks it on a timer, and a matcher exception fails open rather than dropping the request. A bug in blocking should cost an advert, not the web.

10. Read-aloud uses the basic Windows voices

Text-to-speech works offline with no download, through System.Speech and the SAPI5 voices Windows ships. It reads AI replies, the page selection, or a whole article.

The limit is voice quality, and it is worse than it first appears. Windows keeps voices in two separate registries, and System.Speech only reads one of them:

Registry Voices (measured on a Win 11 machine) Reachable today
Speech\Voices — SAPI5 David Desktop, Zira Desktop yes
Speech_OneCore\Voices David, Mark, Zira no

The "Desktop" voices are the oldest and most robotic ones Windows ships. The OneCore set — what Narrator uses — is noticeably better, already installed, and invisible to us.

Installing Windows' natural (neural) voices does not help either. They install into the OneCore registry too, so Settings → Accessibility → Narrator → Add natural voices would leave Nightjar Browser's voice list unchanged. An earlier version of this document claimed they would "appear through the same API with no code change"; that was wrong, and only checking the registries showed it.

Reaching any of them means moving to WinRT Windows.Media.SpeechSynthesis, which enumerates OneCore. The cost is that WinRT returns an audio stream rather than playing it, so playback control — today free from System.Speech — has to be rebuilt around a media player. Pause mid-sentence is the part that gets hard; stop-between-chunks stays easy because the text is already chunked.

Not done because the current voices are adequate for the use case. Piper remains the real quality jump; neither Windows engine reaches it.

Also worth knowing:

Piper (ONNX) remains the upgrade path if the built-in voices prove too grating: onnxruntime is already a dependency, and voices are 25–60 MB each. It was not taken because the Windows engine needs no download at all.

11. Context length varies a lot by model and machine

The context window is sized per model and machine rather than fixed, so the same build behaves very differently across hardware. Measured from the real GGUF headers:

model 6 GB VRAM 8 GB VRAM 12 GB VRAM 16 GB RAM, no GPU
llama-3.2-1b 32k 32k 32k 32k
qwen3.5-2b 32k 32k 32k 32k
qwen3.5-4b 32k 32k 32k 32k
gemma-4-e2b-it 32k 32k 32k 32k
gemma-4-e4b-it 8k (86% GPU) 24k 32k 32k
qwen3.5-9b 8k (81% GPU) 16k 32k 32k
gemma-4-12b-it 8k (62% GPU) 8k (88% GPU) 32k 32k

(16 GB of system RAM throughout; 32k is the automatic ceiling, and a larger window can be chosen in settings where the memory allows.)

What drives the spread is mostly the weights, not the cache: every model in the catalogue keeps a whole-context cache in only a few of its layers. Qwen 3.5 makes three layers in four recurrent, and Gemma 4 limits five in six to a window of 512–1024 tokens, so Gemma 4 12B's cache is 16 KB a token where a plain design of the same size would need several hundred. The planner reads this from each file's header and charges the windowed layers as a fixed cost; the figures were checked against llama.cpp's own allocation.

Consequences: on a 6 GB card the mid-size models run partly on the CPU and slow down accordingly, and --cache-type-k q8_0 (the "longer context on small GPUs" lever) trades a little quality for roughly double the window.

12. Runtime verification is thin

The test suite (783 passing) is unit and integration level. It exercises Python logic and parses ui/sidebar.html as text; it does not drive the real WebView2 control.

Consequences seen in practice: a sidebar reply handler read an undeclared variable and threw on every model reply, rendering the chat silently dead, while the whole suite stayed green. tests/test_sidebar_chat_render.py now guards that specific class of bug, but the general gap remains — a change to sidebar JavaScript is not covered until someone runs the app.


13. Hidden-instruction removal is partial, and is not protection

The assistant reads whatever page is open, so whoever controls a page controls part of the prompt. engine/injection.py and the probe in shell/tab.py take text a person cannot see out of the model's context and report what was removed: CSS-hidden passages the browser confirms are invisible, and the invisible-character families (zero-width, bidi overrides, the Unicode Tag block). Instruction-shaped phrasing in visible text is reported and deliberately left alone.

This raises the cost of the attack. It does not stop it, and nothing in any description of this product should say that it does. What is known to get through:

The thresholds are guesses that no attacker has yet pushed on: passages under 12 characters ignored, 40 findings per page, font sizes between 4px and 8px allowed through, and a colour-distance floor of 24.

TODO.md item 11 tracks the work. The honest description is "hidden text is removed where it can be found" -- never "protected against prompt injection".

Being told is now the exception rather than the rule, changed on 2026-08-28 after the notice was seen firing on ordinary sites. On a web page nothing is said in the conversation at all: the removal happens and a line goes to Activity. Only an attached file speaks up, and only for findings that read as an attempt to steer the answer.

The reasoning, since it reverses an earlier decision. The probe recognises techniques, and the techniques are shared -- clip: rect(0,0,0,0) and left:-9999px are how an accessible site writes "Skip to content" and how an attacker hides a sentence. nytimes.com produced forty-one findings and no attack. A warning above every reply teaches people to dismiss warnings, and the one that matters would have arrived looking like the forty-one before it.

The cost is stated plainly: a page that attacks the assistant now does so without the user being told in the conversation. What protects them is the removal, not the notice. Where the removal fails -- a passage found but not matched in the extracted text, removed: False -- there is now no visible sign of it, and that is the gap this trade opens.


14. Some blank ad space remains, by choice

Blocking an advert stops it loading; the box the page kept for it is a separate problem. Nightjar Browser collapses frames and images whose request was refused and applies each site's own hiding rules, which clears most of it.

It does not, by default, apply the filter lists' general hiding rules -- the ~9,000 like ##.ad-slot that hide anything with an ad-like class name on any site. Sites that check for ad blockers plant bait elements carrying those names, see them hidden, and show a "disable your ad blocker" wall: measured on foxnews.com, the wall appeared with them and not without them. Brave's standard mode makes the same choice.

They are available as Hide more ad space (aggressive) in Privacy & blocking, with that cost stated beside it. On macOS neither the collapsing nor the general rules are wired up yet (see MACOS_HANDOFF.md).