Build the Darth Vader voice assistant, from an empty folder to a working app
You talk or type, Claude writes the reply, and the app speaks it back through a helmet avatar that reacts to every state. This guide covers the tools to install, the three API keys to collect, the prompts to give the Claude Code agent, and the checks that prove each part works.
How a turn moves through the system
- 01 CaptureBrowserMicrophone audio or a typed message
- 02 TranscribeDeepgramSpeech to text, with live captions
- 03 ReplyClaudeStreams the answer sentence by sentence
- 04 SpeakElevenLabsTurns each sentence into audio
- 05 PerformBrowserVoice effect chain and the avatar
The server holds every API key. The browser never sees them. One WebSocket per session carries audio in both directions along with the state events that drive the avatar.
Prerequisites and links
Everything you need to download or sign up for, in one table. Each row is covered in detail in the step named on the right.
| What | Why you need it | Link | Step |
|---|---|---|---|
| Node.js 22 or newer | Runs the server and the web app. This project was built on Node 26. | nodejs.org/en/download | 01 |
| Nimbalyst | Desktop workspace that runs the Claude Code agent beside your files. | nimbalyst.com/download | 02 |
| Claude Code | The coding agent that writes the project. | code.claude.com/docs/en/quickstart | 03 |
| Claude API key | The assistant's replies. | platform.claude.com/settings/keys | 04 |
| Deepgram API key | Speech to text. | console.deepgram.com | 04 |
| ElevenLabs API key and voice ID | Text to speech. | elevenlabs.io/app/developers/api-keys | 04 |
| Chrome, Edge or Safari | Opens the app and grants microphone access. | Already installed on most machines | 07 |
Check your Node version before you start:
node --version # v22 or higher
npm --version
Install Nimbalyst
Nimbalyst is a free, open-source desktop app for macOS, Windows and Linux. It works with your existing Claude Code subscription or API key.
| Platform | Direct download |
|---|---|
| All downloads | https://nimbalyst.com/download/ |
| macOS, Apple Silicon | github.com/Nimbalyst/nimbalyst/releases/latest/download/Nimbalyst-macOS-arm64.dmg |
| macOS, Intel | github.com/Nimbalyst/nimbalyst/releases/latest/download/Nimbalyst-macOS-x64.dmg |
| Windows | github.com/Nimbalyst/nimbalyst/releases/latest/download/Nimbalyst-Windows.exe |
| Linux | github.com/Nimbalyst/nimbalyst/releases/latest/download/Nimbalyst-Linux.AppImage |
| Documentation | https://docs.nimbalyst.com |
- Open the downloaded file. On macOS, drag Nimbalyst into Applications. On Windows, run the installer. On Linux, mark the AppImage as executable and run it.
- Create an empty folder for the project, for example
darth-vader-assistant. - Launch Nimbalyst and open that folder. The window has three panels: files on the left, the editor in the center, and the agent on the right.
Open the Claude Code agent
Install Claude Code once, sign in, then start it from the terminal or from the agent panel in Nimbalyst. Both routes run the same agent.
Install
curl -fsSL https://claude.ai/install.sh | bash
irm https://claude.ai/install.ps1 | iex
brew install --cask claude-code # Homebrew winget install Anthropic.ClaudeCode # WinGet
Open a new terminal window and confirm the install. A working install prints a version number followed by (Claude Code).
claude --version
Sign in and start a session from the terminal
cd ~/repos/darth-vader-assistant claude
- On first run Claude Code asks you to log in. Follow the prompt and finish the sign-in in your browser.
- Use a Claude Pro, Max, Team or Enterprise subscription, or a Claude Console account with prepaid credits.
- To switch accounts later, type
/logininside the session. Type/helpto list commands.
| Command | What it does |
|---|---|
claude | Start an interactive session in the current folder |
claude "task" | Start a session with a first prompt |
claude -c | Continue the most recent conversation in this folder |
claude -r | Pick a previous conversation to resume |
Shift + Tab | Cycle the permission mode inside a session |
Start a session from Nimbalyst
- Open the project folder in Nimbalyst.
- Create a new session in the agent panel on the right.
- Choose a Claude model and a reasoning level in the session composer.
- Type your prompt and send it. The agent reads and edits files in the open folder, and every change appears in the editor as it lands.
Nimbalyst works with your existing Claude Code subscription or API key. If the agent panel asks you to authenticate, run claude once in a terminal and complete the login there.
Get the API keys
Three services, three keys, plus one voice ID. Each key is shown in full only once, at the moment you create it. Copy it straight into your .env file.
Paste keys directly into .env. Do not paste them into a prompt, a screenshot or a commit. Account names, key values and balances are masked in the screenshots below.
- Sign in to the Claude Console at platform.claude.com.
- Open Organization settings, then API keys.
- Select Create key in the top right.
- The console first suggests identity federation. For a project on your own machine, choose Continue with an API key.
- Name the key
vader-assistant, pick an expiry, choose a scope, and select Create key. - Copy the key, which starts with
sk-ant-, and paste it into.envasANTHROPIC_API_KEY. - Make sure the account has credits under Billing. Requests fail without them.
4b. Deepgram API key
Open Deepgram Console- Sign in at console.deepgram.com and select your project in the top left.
- In the left menu under Manage, open API Keys.
- Select Create a New API Key.
- Enter
vader-assistantas the friendly name and choose an expiration. The default role is enough for transcription. - Select Create Key, copy the secret, and paste it into
.envasDEEPGRAM_API_KEY. Deepgram cannot show the secret again.
- Sign in at elevenlabs.io, open Developers at the bottom of the left menu, then the API Keys tab.
- Select Create Key and name it
vader-assistant. - Leave Restrict Key on and set Text to Speech to Access. The app needs nothing else.
- Select Create Key, copy the key, and paste it into
.envasELEVENLABS_API_KEY. - Open Voices. Pick a deep, slow male voice. Open the three-dot menu on its row and choose Copy voice ID.
- Paste the ID into
.envasELEVENLABS_VOICE_ID.
Use a stock voice from the library. The Vader character comes from the effect chain in the browser: a pitch shift, a low-shelf boost, a short comb filter for helmet resonance, and a synthesized breathing loop. Cloning a film voice is not permitted by the provider and is not needed.
Configure the project
Install dependencies and put the keys where the server reads them. The .env file is ignored by git.
npm install cp .env.example .env
Open .env and fill in the values from step 04:
# Claude ANTHROPIC_API_KEY= CLAUDE_MODEL=claude-opus-5-5 CLAUDE_EFFORT=low # Speech to text DEEPGRAM_API_KEY= # Text to speech ELEVENLABS_API_KEY= ELEVENLABS_VOICE_ID= # Server USE_MOCK_PROVIDERS=false PORT=8787 WEB_ORIGIN=http://localhost:5173 DATABASE_PATH=./data/vader.db
| Variable | Required | Default | Purpose |
|---|---|---|---|
ANTHROPIC_API_KEY | Yes | None | Authenticates requests to Claude |
CLAUDE_MODEL | No | claude-opus-5-5 | Model that writes replies. claude-sonnet-5-5 answers faster |
CLAUDE_EFFORT | No | low | Reasoning effort. Low keeps replies quick |
CLAUDE_MAX_TOKENS | No | 2000 | Upper bound on reply length |
DEEPGRAM_API_KEY | Yes | None | Authenticates speech to text |
DEEPGRAM_MODEL | No | nova-3 | Transcription model |
ELEVENLABS_API_KEY | Yes | None | Authenticates text to speech |
ELEVENLABS_VOICE_ID | Yes | None | The stock voice to speak with |
ELEVENLABS_MODEL | No | eleven_flash_v2_5 | Low-latency speech model |
USE_MOCK_PROVIDERS | No | false | Set to true to run the whole loop with no keys and no API calls |
PORT | No | 8787 | Server port |
WEB_ORIGIN | No | http://localhost:5173 | The only origin allowed to connect |
DATABASE_PATH | No | ./data/vader.db | SQLite file for sessions and messages |
Set USE_MOCK_PROVIDERS=true and the full loop runs on built-in mocks. This is the fastest way to confirm the app starts before you spend any credits.
Build it with the agent
Give the agent one phase at a time and review the result before moving on. The prompts below follow the phases this project was built in. Start each one in a session opened on the project folder.
Phase 0. Plan and project rules
Plan a personal voice assistant with an animated Darth Vader style helmet avatar at the center of a dashboard. I talk, it answers out loud, and later it will act on my calendar and mail through MCP connectors. Stack: Vite, React, TypeScript and Tailwind for the web app. Node, Fastify and ws for the server. SQLite for storage. npm workspaces with apps/web, apps/server and packages/protocol. Providers: Claude for replies, Deepgram for streaming speech to text, ElevenLabs for streaming text to speech. Put each behind an interface in apps/server/src/providers so any one can be swapped, and add mock providers so the app runs with no keys. Write the plan to a markdown file with phases, a latency budget, the WebSocket protocol, the data model and the risks. Then write CLAUDE.md with the layout and the rules: keys live in .env and are never printed or committed, and theme assets stay in one swappable folder. Do not write app code yet.
Phase 0. Design system
Propose three visual directions for the dashboard as standalone HTML previews. I want a dark command bridge feel: near-black surfaces, hairline structure, monospace labels, one red accent, and the avatar as the only element that glows. After I pick one, write DESIGN.md with binding tokens for type, color, spacing, radius, shadow, motion and icons, plus a banned list. From then on, no value outside DESIGN.md goes into the app.
Phase 1. App shell and typed chat
Implement phase 1 of the plan. Build the design tokens and UI primitives from DESIGN.md, then the layout: session sidebar, center stage, bottom dock and transcript panel. Add the SQLite schema and migrations, session create, rename, delete and search, the WebSocket gateway, and the shared protocol package. Wire streaming text chat with Claude through the provider interface, with prompt caching and refusal handling. The system prompt is a frozen string written for speech: short plain sentences, no markdown, numbers and dates in spoken form. Done means a typed conversation persists, reloads and resumes. Add tests and run them.
Phase 2. Voice loop
Implement phase 2 of the plan. Capture the microphone in an AudioWorklet and send 16 kHz PCM frames over the WebSocket. Forward them to Deepgram and show interim transcripts as live captions. On end of turn, send the final transcript to Claude. Buffer Claude's stream in a sentence chunker and send each complete sentence to ElevenLabs so speech starts before the full reply exists. Play the returned audio through a PCM player worklet, then through the Vader effect chain: pitch shift down, low-shelf boost, low-pass, short comb filter, compressor. Add a synthesized breathing loop that ducks under speech. Add a microphone picker, a Stop button bound to Esc, and a script that checks each provider with one small real request.
Phase 3. Avatar
Implement phase 3 of the plan. Build the helmet as original procedural geometry in React Three Fiber, with a red key light and bloom inside the avatar canvas only. Drive it from a small state machine with four states: idle, listening, thinking and speaking. Ease between states so nothing snaps. The mouth grille glow and the ring around the avatar follow the audio level from the analyser. Respect prefers-reduced-motion. Fall back to a 2D helmet when WebGL is unavailable.
Review pass
Run typecheck and the full test suite, then run the provider check. Open the app in a browser, send one typed message, and take screenshots of the empty, thinking, speaking and finished states at desktop width and at 390 pixels wide. Report what you verified, what you could not verify, and any failures with their output.
Ask for a plan before code. Build one phase per session. Ask for evidence such as test output and screenshots, and read it before you accept the work.
Run and verify
Two processes: the server and the web app. Start each in its own terminal.
npm run dev:server # http://localhost:8787
npm run dev:web # http://localhost:5173
Open http://localhost:5173. The server takes a few seconds to start. The Providers list in the lower left should show Ready three times.
Check each provider with one real request
npm run check:providers -w @vader/server
OK Language model (Claude): first text in 1338 ms, done in 1393 ms,
served by claude-sonnet-5-5 at low effort, stop reason end_turn
OK Speech to text (Deepgram): connected in 386 ms, model nova-3
OK Text to speech (ElevenLabs): first audio in 308 ms, 1.3 s of audio,
model eleven_flash_v2_5
Run the checks
npm run typecheck npm test
| Package | Test files | Tests | Result |
|---|---|---|---|
| apps/server | 7 | 151 | Passed |
| apps/web | 4 | 21 | Passed |
| packages/protocol | 1 | 10 | Passed |
First conversation
- Select New session.
- Type a message in the dock and press Enter. The state in the top strip moves from Thinking to Speaking, and the reply is captioned under the avatar.
- Press Esc or select Stop to interrupt a reply.
- To speak, select the microphone button and allow access when the browser asks. Use the arrow beside it to pick the input device.
- Press / to search earlier sessions.
The finished app
Captured from the running project on 2026-09-29 with real Claude and ElevenLabs responses. Select any image to view it at full size.
| State | What the avatar does | What you see in the strip |
|---|---|---|
| Idle | Slow breathing motion, faint lens reflections | Grey dot, Idle |
| Listening | Red rim light rises, ring shows microphone level | Red dot, Listening, device name |
| Thinking | Head lowers slightly, slow pulse on the ring | Red dot, Thinking |
| Speaking | Grille glow follows speech level, ring shows output | Red dot, Speaking, response time |
Troubleshooting
The problems most likely to stop a first run, and the fix for each.
| Symptom | Cause | Fix |
|---|---|---|
| A provider does not show Ready | Its key is missing or misspelled in .env | Check the variable name, save the file, and restart the server. Run the provider check to see which one fails |
| Claude check fails with a billing error | The Console account has no credits | Add credits under Billing in the Claude Console |
| Text to speech is skipped | ELEVENLABS_VOICE_ID is empty | Copy a voice ID from the Voices page |
| ElevenLabs returns a permission error | The key is restricted without Text to Speech access | Edit the key and set Text to Speech to Access |
| The page shows Offline | The server is not running, or it is still starting | Start npm run dev:server and wait a few seconds |
| The browser cannot connect after a port change | WEB_ORIGIN no longer matches the web app address | Set WEB_ORIGIN to the exact address in the browser |
| The microphone button does nothing | Microphone permission was denied | Allow the microphone for localhost in the browser's site settings, then reload |
| The wrong microphone is used | A virtual device is the system default | Open the menu beside the microphone button and choose the built-in device |
| The app stays in Thinking | A turn did not complete | A watchdog resets the turn after 20 seconds and shows a notice. Send the message again |
| Replies feel slow | Time to first text from the model | Set CLAUDE_MODEL=claude-sonnet-5-5 and keep CLAUDE_EFFORT=low |
claude is not found after install | The install directory is not on your PATH | Open a new terminal window. If it persists, see the install troubleshooting page in the Claude Code docs |
Status and limits
What has been checked in a browser, what is written but not yet checked by a person speaking and listening, and what is still to build.
- Typed conversation end to end with real Claude and real ElevenLabs speech
- Sessions: create, rename, delete, search
- Transcript with timestamps
- 3D avatar with four states
- Narrow-screen layout
- All three providers pass the provider check
- Microphone capture with a person speaking
- Deepgram transcription of real speech
- The Vader effect chain, judged by ear
- The breathing loop
- Interrupting by voice. Stop and Esc work today
- Push-to-talk
- Voice settings sliders
- Calendar and mail connectors
- Encrypted connector tokens
| Measured on typed turns | First text from Claude | First audio |
|---|---|---|
claude-sonnet-5-5, six turns | 0.9 to 2.3 s | 1.7 to 3.0 s |
| Short reply, fresh session | 1.7 to 2.2 s | 2.2 to 2.9 s |
| Target | Not set | 1.5 s |
The helmet likeness and the character name belong to Lucasfilm and Disney. This is fine as a private project on your own machine. Before you publish or share the app, swap the avatar, name and sounds for original ones. They all live in apps/web/src/theme-assets for that reason.
Checklist
Tick items as you go. Your progress is remembered in this browser.