Bare Docs

Build a local AI chat TUI with QVAC

Layer bare-tui and QVAC on-device inference onto the hello-pear-bare template to build a terminal chat app that runs an LLM on the user's own machine—no API key, no cloud round-trip.

This guide extends hello-pear-bare with a second worker that loads and runs a language model on the user's own machine, and a bare-tui terminal UI on top. The reference implementation is hello-pear-qvac-tui.

╭──────────────────────────────────────────────────────────────────╮
│ ◆ hello-pear-qvac  v0.0.0-rc.0                                   │
│ model LLAMA_3_2_1B_INST_Q4_0   ready                             │
╰──────────────────────────────────────────────────────────────────╯
 LLAMA_3_2_1B_INST_Q4_0 loaded — ask it anything.

 ❯ Name one ocean. Answer in 3 words.

 Pacific Ocean

╭──────────────────────────────────────────────────────────────────╮
│ ❯ ask the model something…                                       │
╰──────────────────────────────────────────────────────────────────╯
  ↵ send · ctrl+t thinking · pgup/pgdn scroll · ctrl+c quit

Ask a question, watch the answer stream back—generated entirely on the machine it's running on. No API key, and it works offline once the model is cached.

This is a delta-only how-to. hello-pear-bare's CLI entry, App/ready-resource shape, and its hello-pear-worker-backed over-the-air updates are covered in Start from the hello-pear-bare template—read that first if you haven't. This guide covers what's added: the inference worker, the bare-tui UI, and QVAC's model/context knobs.

Before you begin

What's added

LayerChange
DependencyAdd @qvac/inference, which registers the llamacpp-completion plugin it ships for GGUF/llama.cpp (backed by the native @qvac/llm-llamacpp peer dependency), plus bare-tui and bare-tui-updater.
A second workerworkers/qvac.js owns the QVAC SDK and the loaded model, alongside workers/main.js—the same hello-pear-worker OTA worker hello-pear-bare already ships.
A client for itlib/inference.js—a ready-resource that runs the worker and turns its frames into events, deliberately the same shape as app.js.
A UIui/app.js (state, keys, layout) and ui/transcript.js (drawing one conversation entry), a plain bare-tui model—your "frontend" is the terminal.

bin.mjs is still the only place everything meets: it constructs both workers' clients and bridges their events into bare-tui messages.

Clone and run

git clone https://github.com/holepunchto/hello-pear-qvac-tui
cd hello-pear-qvac-tui
npm install
npm start

First run downloads the default model (LLAMA_3_2_1B_INST_Q4_0, 0.77 GB—see the table below) and caches it, so it takes a while; later runs start in seconds. Any other model is one flag away:

npm start -- --model QWEN3_1_7B_INST_Q4
keydoes
↵send
escinterrupt the answer being generated
ctrl+tshow/hide a thinking model's reasoning
pgup/pgdnscroll the transcript (mouse wheel too)
ctrl+cinterrupt if busy, otherwise quit

Over-the-air updates work exactly as in hello-pear-bare—see Create a valid upgrade link. Until you set a real pear:// link, bin.mjs here skips constructing the updater rather than throwing, and just runs without it.

Map the template

PathWhat it isDo you edit it?
bin.mjsCLI flags, and the only place both workers meet the UI.Yes—your startup and CLI.
app.jsThe OTA updater client—identical role to hello-pear-bare's app.js.Sometimes—app lifecycle.
workers/main.jsThe hello-pear-worker OTA worker—same package hello-pear-bare ships.Rarely.
lib/inference.jsThe UI-side client for the inference worker.Rarely—it's a thin, stable wrapper.
workers/qvac.jsOwns the QVAC SDK and the loaded model. Registers the plugin, loads the model, runs completions.Yes—this is where you change what's asked and how.
ui/app.jsThe whole UI—state, keys, layout.Yes—change how it looks.
ui/transcript.jsDraws one conversation entry (pure).Sometimes—rendering.

workers/main.js is hello-pear-worker inlined, the same worker every hello-pear-* template ships—see Map the template on the hello-pear-bare page. It has nothing to do with inference; it only carries the OTA updater.

Two workers, one pattern

Loading a GGUF model blocks the thread it runs on for seconds—in-process, the spinner would visibly stutter. So, like the OTA updater, inference gets its own Bare worker, spoken to over a FramedStream. lib/inference.js is deliberately the same shape as app.js: a ready-resource whose _open() spawns the worker with PearRuntime.run and wraps its IPC in a FramedStream. Learning one teaches you the other.

hello-pear-bare's updater client:

app.js
  _open() {
    this.IPC = PearRuntime.run(require.resolve('./workers/main.js'), [
      String(this.updates),
      this.version,
      this.upgrade,
      this.name,
      this.dir,
      this.app || ''
    ])
    this.pipe = new FramedStream(this.IPC)

    this.pipe.on('data', (data) => this._onmessage(data))
    this.pipe.on('error', (err) => this.emit('error', err))
    this.IPC.on('error', (err) => this.emit('error', err))
    this.IPC.on('exit', (code) => {
      if (code === 0 || this.closing !== null || this.closed) return
      this.emit('error', new Error(`Updates worker exited with code ${code}`))
    })
  }

The inference client, same shape:

lib/inference.js
  constructor({ model, ctxSize, gracePeriod, verbose } = {}) {
    super()

    this.model = model || 'LLAMA_3_2_1B_INST_Q4_0'
    this.ctxSize = ctxSize || 8192
    this.gracePeriod = gracePeriod ?? 5000
    this.verbose = verbose === true
    this.loaded = false
    this.percentage = 0

    this.IPC = null
    this.pipe = null

    this._seq = 0

    // Resolves when the worker confirms it has unloaded, or when the grace
    // period runs out — a wedged native unload must not block quitting.
    this._onclosed = null
    this._closed = new Promise((resolve) => {
      this._onclosed = resolve
    })
  }

  _open() {
    this.IPC = PearRuntime.run(require.resolve('../workers/qvac.js'), [
      this.model,
      String(this.ctxSize),
      this.verbose ? '1' : '0'
    ])
    this.pipe = new FramedStream(this.IPC)

    this.pipe.on('data', (data) => this._onmessage(data))
    this.pipe.on('error', (err) => this.emit('error', err))
    this.IPC.on('error', (err) => this.emit('error', err))
    this.IPC.on('exit', (code) => {
      if (code === 0 || this.closing !== null || this.closed) return
      this.emit('error', new Error(`Inference worker exited with code ${code}`))
    })
  }

A resident model holds native handles that keep the process alive, so _close() sends { t: 'close' } and waits (with a grace period) for the worker to confirm it unloaded before destroying the pipe—skip that handshake and the app hangs on exit instead of quitting.

How one question flows

Press enter, and the answer comes back through four layers:

ask Cmd { t:'ask', id, history } sdk.completion(...) contentDelta { t:'delta', id, text } 'delta' event qvac.delta message ui/app.js _submit() lib/inference.js ask() workers/qvac.js ask() @qvac/inference bin.mjs bridge

The id on every frame is how a late token from an interrupted answer gets dropped instead of appended to the next one—ui/app.js tracks the current askId and ignores anything else.

One JSON object per frame, t is the tag. This is the whole protocol between the two, not just the ask/delta pair the diagram above shows—progress and loaded cover the download and startup that happen before the first question, closed the shutdown after the last:

directionframemeaning
→{ t:'ask', id, history }answer this conversation
→{ t:'cancel', id }stop that answer
→{ t:'close' }unload the model and shut down
←{ t:'progress', percentage }model download, 0–100
←{ t:'loaded', model, ctxSize }resident and ready
←{ t:'thinking', id, text }reasoning, from a thinking model
←{ t:'delta', id, text }a token of the answer
←{ t:'end', id, stopReason }finished, cancelled, or cut short
←{ t:'error', id, message }something went wrong
←{ t:'closed' }safe to terminate the thread

workers/qvac.js's ask() produces the outgoing half of that table:

workers/qvac.js
async function ask(id, history) {
  // `completion` returns synchronously; `requestId` is available immediately so
  // a cancel that arrives mid-answer can find this run.
  //
  // captureThinking splits a reasoning model's `<think>` block out of the
  // answer. Without it the tags arrive verbatim in the content — shown to the
  // user, and then fed back as history, where they burn context for nothing.
  const run = sdk.completion({ modelId, history, captureThinking: true })
  inflight.set(id, run.requestId)

  try {
    for await (const event of run.events) {
      if (event.type === 'contentDelta') send({ t: 'delta', id, text: event.text })
      else if (event.type === 'thinkingDelta') send({ t: 'thinking', id, text: event.text })
    }

    // `run.events` ends normally even when the run failed or was cut short at
    // the context limit — the reason only surfaces here. Reporting a bare
    // "finished" after the loop is what makes a truncated answer look complete.
    const final = await run.final
    send({ t: 'end', id, stopReason: final.stopReason || 'eos' })
  } catch (err) {
    // Cancelling rejects `final` with the partial answer attached; that's an
    // outcome, not a failure.
    if (err instanceof sdk.InferenceCancelledError) {
      send({ t: 'end', id, stopReason: 'cancelled' })
    } else {
      send({ t: 'error', id, message: err.message })
    }
  } finally {
    inflight.delete(id)
  }
}

Two things worth keeping if you touch this loop: run.events ends normally even when a completion failed or ran out of context—the reason only surfaces on run.final, which is why ask() awaits it and forwards stopReason rather than declaring victory once the loop exits. And ui/app.js never imports the QVAC SDK—inference reaches it only as the qvac.* messages in the table above, which is why the UI's own tests run in well under a second with no model and no GPU.

The two knobs that matter

Both are CLI flags, parsed in bin.mjs alongside --storage and --no-updates:

bin.mjs
const cmd = command(
  appName,
  summary(pkg.description),
  flag('--version|-v', 'Print the current version'),
  flag('--storage <dir>', 'custom storage directory'),
  flag('--model <name>', 'QVAC model constant to load'),
  flag('--ctx <tokens>', 'context window in tokens (default 8192)'),
  flag('--no-updates', 'disable OTA updates for this run'),
  flag('--verbose', 'log engine and native addon detail to stderr')
)

and resolved, falling back to package.json's qvac field:

bin.mjs
const model = cmd.flags.model || pkg.qvac.model
const ctxSize = Number(cmd.flags.ctx) || pkg.qvac.ctxSize

--model—any model constant @qvac/inference exports. Nothing else in the template changes.

--ctx—everything the model holds at once shares this budget: the conversation so far, its reasoning, and the answer it's writing. The addon's own default is small enough that a couple of turns leave no room to reply and answers stop mid-sentence, which is why this template overrides it:

workers/qvac.js
// The addon's default context window is 1024 tokens, which is small enough
// that a couple of turns of history leave no room for a reply and answers stop
// mid-sentence. Everything the model may hold at once — system prompt, the
// whole conversation, its reasoning, and the answer it is writing — has to fit
// in here, so this is the single most important knob in the template.
const ctxSize = Number(argv(1)) || 8192

If an answer does hit the ceiling, workers/qvac.js reports that in stopReason rather than pretending the model finished (see How one question flows above)—that's what makes the UI able to say so instead of silently truncating.

A handful of models that work out of the box with the llamacpp-completion plugin this template already registers:

ModelSizeGood for
SMOLLM2_360M_INST_Q80.39 GBfastest—smoke tests, weak hardware
LLAMA_3_2_1B_INST_Q4_00.77 GBthe default—a good general baseline
QWEN3_1_7B_INST_Q41.06 GBnoticeably better answers, still fast
QWEN3_4B_INST_Q4_K_M2.50 GBthe sweet spot on most laptops
QWEN3_8B_INST_Q4_K_M5.03 GBthe largest that stays comfortable on a 16 GB machine

Q4 is the everyday quantisation; anything above ~4B on an integrated GPU runs at a few tokens per second. Sizes above are as reported by upstream's README at the pinned commit—not vendored in this repo, so re-check them against the full model table on refresh.

Remix it

Most ideas need no new dependency—just a different prompt and output shape in workers/qvac.js's ask():

BuildChange
Text adventure, NPC dialoguea system prompt; responseFormat for parseable game state instead of prose
An agent that runs thingstools on completion(); execute in the worker and loop
Commit messages, a diff explainerpipe input in as the first history entry, drop the TUI

workers/qvac.js registers exactly one plugin—nothing you don't register is linked in. Swapping the modality (embedding, transcription, translation, image generation, …) means swapping that plugin and installing its peer dependency, which npm won't do for you automatically since QVAC declares peers as optional. The full plugin table, including the non-obvious transitive peers, is in the upstream repo's CLAUDE.md—point a coding agent at that file before asking for changes here; it's dense with the traps this template has already hit and fixed.

Where to go next

Last updated on

Was this helpful?

On this page