Status: v2 scope. Supersedes the v1 interview-copilot PRD (2026-06-22; see git history and 09-MVP-PLAN.md for what v1 shipped). Last updated 2026-07-22 — engine, meeting (Labs), memory, voice/summon, and companion (Labs) are now SHIPPED; per-mode status is inline below. Vision: 00-VISION.md · Delivery: 10-ROADMAP.md.
A cross-platform desktop ambient AI companion (Electron + React + TypeScript). It captures live audio (microphone / system loopback), transcribes in real time, decides when to contribute via a per-mode trigger policy, grounds each contribution in the user’s local documents (and approved memories), and delivers it through an unobtrusive, capture-invisible overlay or its own voice. Interviews (either side), meetings, tutoring, and ambient companionship are modes of one engine.
This is not an offline app. User data is stored locally, but AI features call an AI provider’s API using a key the user provides in Settings — OpenAI today, through a shipped per-capability provider abstraction (§6.7); additional providers are planned.
| Mode | Target user |
|---|---|
| Interview Copilot | candidate in a context where AI assistance is permitted (practice, allowed take-homes, coaching, accessibility) |
| Interviewer Assist | hiring manager / engineer running interviews, wanting better coverage and fairer evaluations |
| Meeting Copilot | anyone in back-to-back calls who wants context, open threads, and action items caught live |
| Tutor | self-learner with material to master (course, codebase, language, cert) |
| Companion | someone who wants an ambient presence with memory while working or gaming |
| Concept | Was (v1) | v2 meaning |
|---|---|---|
| Profile | the candidate | who you are — background documents (resume et al.), notes, style, language; reused everywhere |
| Context Pack | Job | what this is about — a bundle of documents parsed/embedded as a unit; kind: job \| subject \| project \| custom. A job application, a course, a game, a meeting series |
| Session | live interview / mock / sparring | one run of a mode: mode + profile + optional context pack + per-mode settings; transcript, contributions, and report persist to it |
| Mode | implicit (interview only) | a code-defined preset over the engine: sources + trigger policy + persona + grounding scope + surfaces + overlay layout |
| Memory | — | durable facts BrainCue has learned, stored locally, fully user-editable. Shipped review-first: proposals require explicit approval, and only APPROVED memories ever join grounding |
Migration requirements (lossless, automatic): jobs generalizes to context
packs of kind='job'; sessions gain a mode column (existing rows map from
kind: live→interview, mock/sparring→practice); InterviewType becomes
interview/practice-mode scenario config rather than a session-level universal.
No user-visible data loss; v1.5.x databases must open clean.
Microphone, system loopback (unchanged from v1), screen region + clipboard capture (unchanged), and a summon input: global push-to-talk / typed ask addressed directly to BrainCue from any mode.
OpenAI Realtime GA streaming STT with server VAD (unchanged); chunked STT
fallback. Speaker attribution stays best-effort (interviewer/user labels
generalize to them/you).
A mode declares when the agent contributes:
| Policy | Fires when | Used by |
|---|---|---|
| reactive | a question/request directed at the user is detected (v1’s classifier) | Interview Copilot |
| proactive | salience detected: unanswered question, action item, claim needing context, coverage gap | Meeting Copilot, Interviewer Assist |
| dialogue | it is the agent’s turn in a two-way conversation | Practice, Tutor, Companion |
| summoned | the user explicitly asks (hotkey / push-to-talk / Ask box) | every mode |
Requirements: per-mode sensitivity control, interjection cooldowns, and a hard mute (pause AI) that always wins. Balanced is the default posture for meetings (since 2026-09-07 — quiet shipped first and users reported “question detection is not working”): a question asked in the room streams a grounded answer; at quiet it becomes an open-question card instead. Questions bypass the card cooldowns at every level; everything else still obeys them.
Retrieval over profile + the session’s context pack + approved memories, exactly as v1’s RAG path. The “data sent to OpenAI” transparency panel remains a hard requirement in every mode.
Per-mode persona prompt; streaming; the never-invent rule and risk warnings carry over from v1 §7.3 unchanged.
The engine talks to capabilities, not to OpenAI: a Provider interface exposes
chat (streaming), embeddings, stt (realtime + chunked), tts/speech,
and vision, and each concrete provider declares which it implements.
OpenAI remains the reference implementation and default. The seam is CUT and
shipped (src/main/providers/registry.ts resolves per capability); no second
provider is registered yet — Settings → Providers surfaces the layer with the
planned providers marked “Coming soon” rather than offering a dead choice.
Everything in the v1 PRD §7 (profiles, jobs→context packs, documents, RAG, live session, overlay, coding/screenshot mode, privacy, tour) remains in force verbatim. This mode is the regression gate: no v2 refactor may degrade it.
Inputs: your role’s JD (context pack) + the candidate’s resume (document).
Live: suggested opening questions, follow-up suggestions generated from the
candidate’s last answer (reuses interviewer.ts), and a coverage tracker
(which competency areas have/haven’t been probed). Post-session: a structured
evaluation draft (reuses feedback.ts). Same overlay, opposite chair.
Balanced by default. A question asked in the room is answered right away, in
the Cue Card, grounded the same way a summon is; at quiet presence it becomes
an open-question card that quotes it. Proactive contribution cards: relevant
context from the pack (“this was decided in the attached doc”) and action items
/ decisions as they’re spoken. (The v2.0 “hold the question, surface it if two
turns pass without an answer” tracker was removed on 2026-09-07: in real
meetings almost any next turn shared a word with the question, so nothing ever
surfaced.) End of session: meeting summary report (decisions, actions, open
threads). Sensitivity dial from “only when summoned” to “eager”.
As built: deterministic heuristics filter small talk before the salience
classifier ever runs; the Presence dial (summoned/quiet/balanced/active) maps
to explicit confidence floors + cooldowns; gated by its acceptance suite
(meeting.acceptance.test.ts) and surfaced with a Labs badge.
Any context pack of kind subject (textbook chapter, codebase docs, language
notes). Teach / quiz / drill loop — a generalization of v1 sparring: agent
speaks (voice), user answers by voice, per-answer coaching persists to Reports.
Phase 3’s Realtime speech-to-speech makes it a natural conversation rather than
turn-based MP3 exchanges.
Explicitly started, memory-backed ambient presence. A presence dial
(silent observer ↔ chatty) controls the interjection policy. Game-buddy is
Companion + the existing screen-region Vision path pointed at the game (still
open — the one unshipped piece).
As built: CompanionPresence off/on_demand/assistive/proactive (off is a
hard mute), the InterjectionPolicy gate chain in cost order (mute → presence →
DND windows → budget → heuristics → cooldowns → classifier → confidence/
relevance floors → rate cap → dedupe — the LLM only scores, code decides),
approved-only memory recall re-gated on real vector score, per-session hard
budget with a live cost meter in the Cue Card, and quiet hours. Gated by its
scripted evaluation harness (companion.eval.test.ts).
v1’s mock + sparring continue unchanged, re-labelled as the Practice mode
family; they migrate onto the engine’s dialogue policy when Tutor is built
(shared loop), not before.
| Risk | Mitigation |
|---|---|
| Interjection annoyance kills trust in ambient modes | silence-first defaults, sensitivity dial, cooldowns, hard mute |
| Always-on API cost surprises | VAD gating, visible cost estimates, budgets; local STT evaluation |
| Mode sprawl forks the codebase | engine-first rule enforced in review; modes are config |
| Refactor breaks the shipped interview product | parity gate: full test suite + privacy hard test at every phase boundary |
| Rebrand confuses existing users | flows unchanged; tour + changelog explain the widening, not a pivot away |
| Provider lock-in / API drift (Realtime GA is OpenAI-shaped) | provider layer (§6.7): capability interfaces, per-capability fallbacks, OpenAI as reference impl |