FREE · PRIVATE · LOCAL INFERENCE

RunanycompatibleopenLLMonyourmachine.AnAIassistantthatneversendsyourdataout.

Download any supported GGUF model that fits your hardware. Browse normally, but keep your AI prompts, uploaded documents, and page analysis contained locally — zero AI APIs, zero cloud endpoints.

› Stated exactly, because the distinction is the product: this is still a browser, so the sites you visit see your requests as they would anywhere, and non-URL text typed in the address bar goes to Google as a search. What never leaves is what the assistant works on — page content, chat and attached files — unless you point it at an AI engine on another machine yourself, which the model picker marks as leaving it.

$ get nightjar browsermacOS previewInstall stepsWindows 10 (1903+) or 11, 64-bit · macOS 26 on Apple Silicon · no Linux build
nightjar — new tab
ask>

READING
no document
MODEL
Qwen 3.5 9B
INFERENCE
localhost
WHERE THE PAGE GOES

The competition is not
Chrome. It is the sidebar.

An AI sidebar is most useful when it sees the most, and in every other browser everything it sees leaves your computer. That is the structural problem of the whole category, not a setting any of them can flip. Nightjar Browser answers from a model running beside the browser, so there is no endpoint the page could be sent to.

Edge + Copilot→ Microsoft
Chrome + Gemini→ Google
Perplexity Comet→ Perplexity
Arc Max→ OpenAI
Nightjar Browser→ localhost

This matters most where it is not a preference but a requirement: client files, patient records, privileged documents, anything under a policy that forbids sending data to a third party. A smaller model that never leaves the building is a trade some people cannot make and others must.

api.provider.com
CHILD PROCESSllama-server127.0.0.1no api key
PAGE · CHAT · FILES — ON-DEVICE
BrowserSends usage data home by defaultBlocks trackers by defaultSends the open page to a cloud AIKeeps a record of your browsing
ChromeYes (Google)No — Safe Browsing screens malicious sites, not trackersYes, if Gemini features are usedYes — history, cache and cookies, synced if you sign in
EdgeYes (Microsoft)Partial — Tracking Prevention, Balanced by defaultYes, if Copilot is usedYes — history, cache and cookies
Nightjar BrowserNoYes — EasyList and EasyPrivacy, on for everyoneNo — there is no cloud AI in this browserNo — private by default: history, cache and cookies stay in memory

That No is a claim about the code, not a policy. There is no telemetry client. Two things reach the network without being asked each time, both listed in the privacy policy and both switchable in Privacy & blocking: one check for a newer version at launch, which sends the version you have and nothing else, and the ad-blocking filter lists, downloaded from easylist.to every three days. A crash report is a file you read and choose to email. Edge Tracking Prevention is real and worth naming honestly; what it does not do is block ads.

WHAT IT DOES

One guarantee,
and a browser around it.

The assistant is the reason to switch. Everything below it — blocking, banner suppression, tab suspension — is table stakes that keeps the browser competitive, and none of it is a reason to switch on its own: a Chrome user running uBlock Origin already has most of it.

PAGE Q&A

Ask about the page you are on

The sidebar reads the current tab — article, documentation or a PDF — and answers from it, citing which part it used. Replies stream as they generate. A source small enough to fit is sent whole and in reading order rather than as reshuffled fragments.

▸ first words at 0.28 s
ATTACHMENTS

Ask about your own files

Drop a PDF, Excel workbook, CSV or source file into the chat and it is indexed alongside the page. Ask for a total, an average or the highest row of a spreadsheet and Nightjar calculates it over every row, showing its working, instead of trusting a small model's arithmetic.

▸ 49 file types accepted
FIELD ASSIST

Fill the form in front of you

A Fill button appears when a page has fields it recognises, matching them against a profile stored encrypted on this machine. An assistant with reach into credential and payment fields is exactly the one you want running locally — that argument needs no privacy lecture.

▸ profile never sent anywhere
FILTERING

Block ads and trackers

Every request a page makes is checked, inside the browser, against EasyList and EasyPrivacy with Brave's adblock-rust engine, before it is sent. Refused ad frames and images collapse, so no blank box is left behind — and sites that ask you to switch your ad blocker off are not given a reason to. Nothing is decrypted or proxied.

▸ ~141,000 rules, on for everyone
CONSENT

Cookie banners, gone

OneTrust, Didomi, Cookiebot, Quantcast and TrustArc banners collapse on DOM creation; anything unrecognised gets Reject All clicked for you. Sec-GPC: 1 and DNT: 1 go out with it, which several US state laws require a site to honour as an opt-out.

▸ Chrome does not send GPC
MEMORY

Background tabs stop costing anything

45 seconds after a tab loses focus its renderer is frozen and its memory handed back; after 30 minutes unseen, or sooner if memory runs short, it is unloaded. Fifty tabs of Wikipedia and Hacker News: 6,253 MB all live, 1,070 MB with forty-nine frozen, 633 MB with them unloaded.

▸ 90% returned once unloaded
SPEECH

Read pages aloud, offline

Through the voices Windows already ships, so nothing is sent to a speech service to be spoken back to you.

▸ no network needed
SIZING

It picks the model for your machine

Nightjar Browser sizes the model and its context from each GGUF header against your RAM and VRAM, hides models that will not run, and never asks you to choose a layer count. Models you already have in Ollama or LM Studio appear in the same picker.

▸ 7 models in the catalogue
MEASURED, NOT ESTIMATED

Two different claims, worth separating.

Pages load faster with blocking on, mostly by not stalling. The AI is fast because it is local. Only one of those is about the browser, so they are measured separately — one machine, one network, three passes per arm. Indicative of the effect, not a benchmark suite.

The packaged build ships Vulkan

Vulkan runs on NVIDIA, AMD and Intel alike, rather than a separate 1 GB CUDA build for NVIDIA only — and it is a 20x smaller download, 32 MB against 611 MB. Measured on an RTX 3060 12 GB with tests/benchmark_backends.py:

WorkloadCUDAVulkan
Qwen 7B generation63.6 tok/s53.1 tok/s (−16%)
Qwen 1.5B generation187.9 tok/s195.7 tok/s (+4%)
Reading a 10,000-token page5.0 s5.6 s (−10%)

Sixteen per cent on the largest model, and nothing at all on the smallest. These were taken at a different context size from the tables above, so treat the two as separate measurements rather than one series.

WHAT YOU NEED

The browser runs on anything.
The AI panel is what needs hardware.

Because the model runs on your computer instead of someone else’s server. Also required: .NET Framework 4.7.2+ and the WebView2 runtime, both standard on Windows 10 1903 and later. An internet connection is needed for first-time setup; everything after that runs offline.

MinimumRecommendedBest
OSWindows 10 (1903+), 64-bitWindows 11Windows 11
RAM4 GB16 GB32 GB
GPUnone (CPU only)8 GB VRAM12 GB+ VRAM
Free disk2 GB10 GB30 GB+
Models available2 of 77 of 77 of 7, longest context

RAM only — no graphics card

Computed from each model’s own GGUF header, not estimated.

Your RAMAvailableRecommendsContext
2 GBnone——
4 GB2 of 7Qwen 3.5 2B (1.2 GB)32k
8 GB4 of 7Qwen 3.5 2B (1.2 GB)32k
16 GB7 of 7Qwen 3.5 2B (1.2 GB)32k
32 GB7 of 7Qwen 3.5 2B (1.2 GB)32k

The largest model that fits is rarely the best choice without a GPU: a 6.6 GB model on CPU generates a couple of words per second. The recommendation accounts for that, and for whether a model would be squeezed below a useful context.

With a graphics card

For Qwen 3.5 9B (5.3 GB), a typical mid-size model. Nightjar Browser sizes layers and context to your card automatically.

Your VRAMQwen 9B runsLayersContext
4 GBpartly on the GPU16 of 328k
6 GBpartly on the GPU26 of 328k
8 GBentirely on the GPU32 of 3216k
12 GBentirely on the GPU32 of 3232k — every model fits
16 GB+entirely on the GPU32 of 3232k on the largest models too

The 8 GB row shows the trade-off: the whole model runs on the card, and the context is held to 16k to keep it there. Keeping every layer on the GPU is worth far more than a longer context, because a layer moved to the CPU costs speed on every single word. NVIDIA uses CUDA; AMD and Intel use Vulkan.

BEING STRAIGHT ABOUT THE LIMITS

What it cannot do, stated up front.

These are deliberate current boundaries, not bugs to be filed. The full list, with what each gap would take to lift, is maintained honestly in the repository and is worth reading before relying on anything here.

01

Windows first

The Windows build is the one people use. A macOS preview for Apple Silicon (macOS 26 or newer) ships with each release; it runs, but has had far less use, and is not yet signed, so installing it takes one Terminal command. There is no Linux build.

02

A local 7B is not GPT-5

You are trading capability for the guarantee that nothing leaves. That is the right trade for some work and the wrong one for other work.

03

Retrieval matches on words, not meaning

Asking about earnings may miss a page that only ever says revenue. The model is told the page may have nothing to do with the question and to say so rather than stretch it.

04

Small models get comparisons wrong, and invent figures

Measured, not assumed: on the built-in 16-question suite Llama 3.2 1B named the wrong region for highest revenue and supplied a figure for something the page did not contain, and Qwen 3.5 9B got a sum wrong by three million. Ask a small model which row is highest, or to total two values, and check the answer against the page. The browser runs this suite on whatever model you load and tells you where it is weak.

05

The model never sees a whole long document

Only the parts retrieval selected. A summary of a 200-page PDF summarises those parts.

06

No OCR — scanned PDFs yield nothing

Text extraction reads a PDF existing text layer. A PDF that is a picture of a page is refused. Probed and feasible offline through the OCR engine Windows 11 already ships; deferred by decision, not blocked.

07

No image input, even on multimodal models

Attaching an image is refused. Text-only weights cannot take images at all; for multimodal weights the model can see but the app cannot yet send.

$ read LIMITATIONS.md in full