JB - Jot Brief — documentation
A Windows app that transcribes meetings (Google Meet, Teams, Zoom) without a bot, separates and names who is speaking, and takes the conversation to Claude (chat inside the app, a claude.ai project, or the Claude app for Windows). Everything audio-related runs on your computer; only the text goes to Claude, when you ask for it.
This documentation describes the app as it is today. To install it and use it day to day, also see the
README.md.
1. Overview
| What it does | How |
|---|---|
| Records the meeting without a bot | Captures the audio that comes out of your speakers/headset (WASAPI loopback) + your microphone |
| Transcribes locally | Whisper (faster-whisper) on your NVIDIA GPU; without one, it uses the CPU with a smaller model |
| Shows the conversation live | Chat-style bubbles that follow the audio; you on the right, the meeting on the left |
| Separates and names people | Voice separation + voice recognition + names seen in Meet + calendar invitees |
| Chats with Claude during the meeting | Chat column with ready-made buttons (summary, tasks, questions to ask…) |
| Takes it to Claude | Browser, Windows app, or MCP server (Claude Desktop reads the meetings on its own) |
Costs: transcription, voice separation, and recognition are free (they run locally). What uses the Anthropic API
(with your ANTHROPIC_API_KEY) is the panel chat and the meeting's automatic subject. The app shows token usage
and the estimated cost in USD and in BRL (R$).
2. Installation and first use
Requirements: Windows 11, Python 3.12 (managed by uv), and, for speed, an NVIDIA GPU.
uv sync --extra cuda # with an NVIDIA GPU; without a GPU: uv sync
copy .env.example .env # fill in ANTHROPIC_API_KEY
uv run jotbrief setup # downloads the transcription models (~1.6 GB)
uv run jotbrief gui # opens the window
- The API key lives only in the
.envfile (it never goes to Git or onto the screen). - The voice separation models (≈ 32 MB) are downloaded the first time you use 👥.
- Audio devices:
uv run jotbrief deviceslists them; by default the app uses the Windows defaults.
First recording: open the app → choose the language (Português, Inglês, or Automático — Portuguese, English, or Automatic) → click the round record button → speak/listen to the meeting → click again to stop. The meeting shows up in the list on the left.
3. The window
┌──────────────┬───────────────────────────────┬──────────────────┐
│ Meeting │ ● 00:00:00 + Nova reunião │ Claude │
│ list │ [button bar] │ model · effort │
│ (by day) │ Gravando em: title │ usage (US$/R$) │
│ │ audio player │ [ready buttons] │
│ │ transcript bubbles │ conversation │
│ │ │ [question ↑] │
├──────────────┴───────────────────────────────┴──────────────────┤
│ LGPD notice │
└──────────────────────────────────────────────────────────────────┘
3.1 Meeting list (left)
- Each meeting is a card:
YYYY-MM-DD | hh:mm | Subject(date and time in bold), number of utterances and, if a task is in progress, the line⏳ identificando falantes…(identifying speakers…). - Grouped by day; the day button (
▾ HOJE · 2026-10-02 (7), where HOJE = TODAY) collapses/expands it. When you open the app, only the latest day with a transcript is open. The selected meeting, or the one being recorded, is never hidden. - Buscar reuniões (Search meetings): searches the title and the text of the transcripts.
- 🗑 moves the meeting to the Windows Recycle Bin (you can recover it). It won't let you delete the one being recorded.
- + Nova reunião (New meeting) leaves none selected: the next recording creates a new one (a "Nova reunião" row appears at the top). Recording with a meeting selected continues it (it becomes "Parte 2" (Part 2), with a separator).
3.2 Button bar (center)
| Group | Buttons |
|---|---|
| Record (Gravar) | Record/stop + time · Nova reunião (New meeting) · language |
| Claude | ✳🌐 browser · ⚙ project · ✳🪟 Windows app (manual) |
| People | 👥 identify speakers · ✏ assign names |
| Reprocess | 🔄 transcribes the audio again |
| Folder | 📁 opens the meeting folder |
| Copy | 📋 copies the transcript in the ideal format for AI |
| Chat | 💬 shows/hides the Claude column |
| Visual | 🎈 balloon mode · light/dark theme |
3.3 Bubbles and player
- You ("Eu" = Me) on the right, in green; the meeting on the left, one color per person. Partial text appears dotted and becomes final when the utterance closes.
- Click the name on a bubble to rename the person (this also teaches the app their voice).
- Player: play/pause, ±10 s, position bar, volume. The text follows the audio like song lyrics; "▶ Ouvir parte" (Play section) plays just one section. The ↓ button jumps to the end; during recording the screen scrolls on its own.
- Balloon mode (🎈): when recording, the window turns into a floating bar with the latest utterance.
3.4 Chat with Claude (right)
- Ready-made buttons: 📝 Resumo detalhado (Detailed summary) · ⚡ Resumo rápido (Quick summary) · 🎯 Pontos-chave (Key points) · ✅ Decisões e tarefas (Decisions and tasks) · ❓ Em aberto (Open items) · 💡 Perguntas para fazer (Questions to ask) · 🔎 Última fala (Last utterance). The text of the request appears when you hover over the button.
- Free conversation: type and press Enter (Shift+Enter breaks the line). The answer arrives gradually; ■ interrupts it.
- During recording Claude reads the transcript up to that moment (refreshed with every question).
- Model (Sonnet 5.5, Opus 5.5, or Haiku 4.5) and effort (low → max) are chosen at the top; they apply to the whole app (chat and subject). Haiku has no effort setting.
- Usage: last answer and meeting total (input tokens, the part that came from cache, output), estimated cost in USD and BRL (R$) (current exchange rate, refreshed every hour; the source and time are shown in the panel).
- Each meeting has its own conversation, saved in
chat.jsonin its folder. 🗑 clears it.
4. How each feature works
4.1 Capture and transcription
- Two channels recorded into a stereo 16 kHz
audio.wav: left = your microphone, right = the meeting audio. - A speech detector (silero-vad, running via ONNX, without PyTorch) cuts the audio into chunks; each chunk goes to Whisper. There are filters against "hallucinations" (ghost subtitles such as "Legenda Adriana Zanotto" (a Portuguese subtitle credit), text in another alphabet, repeated phrases).
- Language: Português, Inglês, or Automático (chooses between pt and en per chunk).
4.2 Microphone echo (without a headset)
If the microphone "hears" the speaker, the meeting sound would come back as if it were you. The app handles this in layers: 1. Text: discards mic utterances that repeat what the meeting said at the same moment. 2. Audio, live and after stopping: compares the volume of the two channels over time. It is leakage when the microphone is extremely quiet (≤ 0.002 measured) while the meeting is speaking; your voice, even on top of the meeting, is 10–40× louder and is never hidden. 3. Recovery: when two people talk at once, the meeting channel only transcribes the stronger voice; the other may have been left only in the mic leakage. If Meet confirms 2+ people speaking together and the text does not exist in the meeting channel, the chunk comes back as an utterance (with the name of whoever talked over the other).
Limit: with a headset the mic doesn't pick up the meeting; so there is no echo, but a second voice talking over the first is not recorded. Only capturing each participant's audio separately would solve this (not implemented).
4.3 Identify speakers (👥)
- Separates the voices in the meeting channel into Pessoa 1, 2… (Person 1, 2…) (local models: pyannote segmentation + voice embeddings via sherpa-onnx), including splitting an utterance in which two people alternate.
- Runs automatically when you stop the recording, in a separate process (the window doesn't freeze) — ~1 min for 9 min of audio.
- When you click it manually you can enter how many people are speaking (not counting you); this separates similar voices better.
- Settings in
config.toml:diarization_threshold(lower = separates more; 0.3 default) andnum_speakers.
4.4 Assigning names to people
Four sources, from weakest to strongest (what you type is never overwritten):
1. Voice recognition: when you name someone, the app stores their voice print and, in future meetings, assigns the name on its own
(voice_match_threshold 0.75 and voice_match_margin 0.05). It is biometric data: it stays only in %APPDATA%\jotbrief\voices.json
(delete the file to forget everyone).
2. Calendar invitees: with calendar_ics_url (the secret iCal address of Google Calendar), the app finds the event at that time
(including weekly/daily meetings) and puts the invitees at the top of the name list.
3. Names seen in Meet (Chrome extension, below): matches who Meet showed speaking alone with each voice.
4. You: by clicking the name on the bubble or ✏ (a list of people already registered + typing a new one).
4.5 Chrome extension (Google Meet)
extension/ folder — install at chrome://extensions → Developer mode → Load unpacked → extension folder.
- It watches the Meet page and tells the app who is speaking. It shows a badge in the bottom-left corner:
[JB - Jot Brief] Name · app gravando ✓ (app recording ✓; it also says whether the app is closed or not recording).
- Extension icon: the green JB; it gets a red circle when the app is recording.
- Live: when only one person is speaking, the bubble already comes out with their name. After you stop, the voice matching redoes the names and
learns the voice.
- Security: it only talks to 127.0.0.1:47821; the app refuses requests coming from web pages (it requires a header that only the extension
sends).
- The notices are kept in falantes_meet.jsonl in the meeting folder.
- It is experimental: Meet doesn't expose "who is speaking" in a stable way, so detection is based on interface activity and may need
adjusting if Meet changes (DEBUG in content.js keeps the badge visible).
4.6 Reprocess (🔄)
Transcribes audio.wav again with the current language and filters (backing up the previous transcript to transcricao.jsonl.bak),
optionally identifying the speakers right away. Useful after improvements to the app or to fix the language.
4.7 Taking it to Claude
- ✳🌐 Browser: opens a new conversation inside your project on claude.ai with the skill's request + the transcript. With
auto_send = truethe app waits for the page to load and presses Enter (only if the front window is a new Claude page). Long calls (> ~7,500 characters in the link) carry the request in the link and the transcript stays on the clipboard for you to paste. - ✳🪟 Windows app (manual): the Claude app only accepts text in new conversations outside projects. So the button opens the project and copies the text: click "Nova sessão" (New session) and paste (Ctrl+V).
- ⚙ Project: sets the project URL (
https://claude.ai/project/…). jb-jot-brief-transcricaoskill:uv run jotbrief skillgenerates the.zip(and thepainel.html) to upload to Claude. The skill asks what to generate and delivers the suggested conversation name and copy buttons.- MCP server:
uv run jotbrief mcp-install(with Claude Desktop closed) registers a read-only server withlist_meetingsandget_transcript; Claude Desktop then reads the meetings on its own. - Copy (📋): one utterance per line, with real date and time: ``` Reunião: 2026-09-29 | 15:52 | Assunto Duração: 00:00:22 | Participantes: Pessoa 1, Eu
2026-09-29 15:52:02 | Pessoa 1 | Conta, Carlos. Por que você está tão animado? ```
4.8 Automatic subject
When you stop, Claude generates a short title (up to 8 words) and the list starts showing YYYY-MM-DD | hh:mm | Subject. Without a key/network, it stays as
"Sem assunto" (No subject) (no harm done).
5. Files
5.1 Per meeting — reunioes\YYYY-MM-DD_HHMM\
| File | Contents |
|---|---|
audio.wav |
stereo 16 kHz (L = mic, R = meeting) |
transcricao.jsonl |
one utterance per line: t0, t1, source (mic/loop), text, speaker, words, echo, recovered |
transcricao.md |
readable transcript |
meta.json |
subject, parts, names (names), which ones were automatic (auto_names), calendar invitees |
falantes_meet.jsonl |
extension notices (second, name, started/stopped) |
chat.json |
conversation with Claude and the usage of each answer |
transcricao.jsonl.bak |
backup before reprocessing/identifying |
app.log |
recording log |
5.2 App-level — %APPDATA%\jotbrief\
config.toml · people.json (registered names) · voices.json (voice prints) · prompts.json (the Claude button's requests) ·
cotacao.json (dollar rate in BRL) · models\ (voice separation models).
5.3 Configuration (config.toml)
| Key | Default | What for |
|---|---|---|
language |
pt |
transcription language (pt, en, auto) |
device |
auto |
auto, cuda, or cpu |
model_gpu / model_cpu |
large-v3-turbo / small |
Whisper models |
output_dir |
reunioes |
where the meetings are stored |
claude_model |
claude-sonnet-5-5 |
model for the chat and the subject |
chat_effort |
medium |
Claude's effort in the chat (low…max) |
chat_open |
true |
chat column visible |
claude_project_url |
— | claude.ai project |
skill_name |
jb-jot-brief-transcricao |
skill name (goes at the start of the message) |
auto_send |
true |
presses Enter on its own in the browser |
ask_in_app |
false |
true: the button asks what to generate in the app (otherwise the skill asks in Claude) |
calendar_ics_url |
empty | secret iCal address of Google Calendar |
diarization_threshold |
0.3 |
voice separation (lower = separates more) |
num_speakers |
0 |
number of people, if you know it (0 = automatic) |
voice_match_threshold / voice_match_margin |
0.75 / 0.05 |
strictness of voice recognition |
speaker_threshold |
0.55 |
separation without the voice-change model |
silence_ms / max_segment_s |
600 / 15 |
utterance cutting |
mic_device / loopback_device |
— | device indexes (jotbrief devices) |
theme / volume / floating |
claro / 100 / false |
appearance and balloon mode (claro = light) |
6. Command line
uv run jotbrief <command>
| Command | What it does |
|---|---|
gui |
opens the window |
run |
records and transcribes in the terminal (Ctrl+C to stop) |
devices |
lists audio devices |
setup |
downloads/loads the models |
identify <folder> [--people N] |
separates the voices of a meeting (used by the app, in a separate process) |
learn-voices |
learns the voices of people already named in the saved meetings |
subjects |
generates the subject for meetings with "Sem assunto" |
check-key |
tests the ANTHROPIC_API_KEY |
skill |
generates the Claude skill (skills\) |
mcp / mcp-config / mcp-install |
MCP server; shows the configuration; installs it in Claude Desktop (close it first) |
7. Architecture (for anyone who will tinker with it)
audio.py ─ Track/StereoRecorder ─► session.py ─► transcriber.py (Whisper) ─► transcricao.jsonl
vad.py (silero) │ ├─ echo.py (audio echo, recovery)
│ └─ meet_names.py / meet_bridge.py (extension)
speakers.py (separation) voices.py (recognition) people.py agenda.py (invitees)
reprocess.py jobs.py (identify in a separate process)
chat.py (Claude: streaming, prices, usage) fx.py (dollar→BRL) summarize.py (subject)
claude_link.py prompts.py skill.py auto_send.py claude_config.py mcp_server.py meetings.py
ui.py (PySide6) ui_helpers.py (Qt-free parts) extension/ (Chrome)
- UI golden rule: workers → UI only via
Signal. - Heavy work outside the window: voice separation runs in another process (
jobs.run_identify), because it was native code that froze the interface. - Echo:
echo.EnvelopeLogstores the volume of both channels live;echo.mark_echo_by_audiorepeats the measurement on the full audio (when identifying speakers and when reprocessing). - Time: the session timeline (
session.now()) is the clock for everything: utterances, Meet notices, and position in the audio; when resuming a meeting, it adds what has already been recorded. - Tests:
uv run pytest -q(≈ 95 tests: pure logic, UI without a screen (offscreen), extension, formats). The ones that download models are markedslow.
8. Privacy and security
- LGPD (Brazil's data protection law): recording and transcribing meetings requires the participants' consent (the notice is in the footer). Tell people before you start.
- Audio and voice stay on your PC. Voice prints are biometric data: only in
voices.json, never sent anywhere. - What goes to Anthropic: the transcript text (and the request), only when you use the chat, generate the subject, or send it to Claude.
- API key: only in
.env; never in the code, in Git, or in messages. - Chrome extension: talks only to
127.0.0.1; the app refuses requests from websites. - Calendar: the secret iCal address is equivalent to a read-only password for your calendar; keep it only in
config.toml(if it leaks, use "Reset" in Google Calendar). - Deletion: goes to the Windows Recycle Bin.
9. Known issues and limits
| Situation | What happens / what to do |
|---|---|
| Two people speak at the same time | The meeting channel transcribes the stronger voice. The other only shows up if it leaked through the mic (no headset) and Meet confirmed the simultaneous speech |
| No headset | Some echo may remain; the app filters it by volume. A headset is ideal |
| Many "ghost" people in the separation | Enter the number of people in 👥 or raise diarization_threshold |
| Wrong voice name | Fix it with ✏; the manual correction carries more weight and is never overwritten |
| Claude app for Windows | Doesn't accept text + project: the button is manual (it opens the project, you paste) |
| Extension doesn't show the right name | Reload the extension and the Meet tab; check the [JB - Jot Brief] badge and the Console (F12) with [JotBrief] |
| I can't delete a meeting | You can't while it is being recorded; stop the recording first |
| Chat doesn't answer | Check the message in the panel: usually a missing/invalid key or no internet (uv run jotbrief check-key) |
| Chat cost | Each question resends the transcript as context (with cache). Long meetings cost more per question |
| Teams/Zoom desktop | Audio capture works for any app; the automatic names from Meet only apply to Meet in the browser |
10. Glossary
- Loopback: audio that comes out of the computer (the voice of the others in the meeting).
- Pessoa N (Person N): automatic label for a separated voice.
- Echo/leakage: the microphone picking up the speaker.
- Effort: how much Claude "thinks" before answering (more effort = slower and more expensive).
- MCP: the protocol through which Claude Desktop reads the meetings as a tool.
- Skill: a package of instructions you install in Claude to handle JB transcripts.