IntervAI
The mock interviewer that remembers what trips you up — and won't let you ignore it.
About
An AI-powered mock interview platform that actually remembers you. Most interview tools forget you the moment you close the tab — IntervAI tracks your weak areas across sessions and makes sure the next interview drills you harder on exactly what tripped you up. It conducts role and company-specific interviews (Google, Amazon, Startup), scores every answer using the STAR framework, and builds a persistent memory of your gaps over time. The result is a compound prep loop: the more you practice, the smarter and more targeted the questions become.
Core Features
Memory System
Tracks weak topics across sessions. Next interview injects those gaps into Claude's prompt — 40% of questions target your actual blind spots.
STAR Scoring
Claude grades each answer on Situation, Task, Action, Result (0–2.5 each) with line-by-line feedback on exactly what was missing.
Voice Mode
Speak your answers. The AI interviewer responds in ElevenLabs voice. No keyboard — full real-interview simulation.
Company Mode
Google L4, Amazon (Leadership Principles), FAANG, or Startup — question style and culture pressure matched per company.
JD Predictor
Paste any job description → Claude predicts the top 10 questions for that specific role and company before you even apply.
Performance Heatmap
GitHub-style 365-day grid of practice intensity + a topic coverage map that makes your preparation blind spots impossible to ignore.
Voice AI Pipeline
End-to-end flow from microphone input to AI interviewer voice response
Web Speech API vs Whisper — why we chose the browser
| Metric | Web Speech APIused | Whisper API |
|---|---|---|
| Cost | Free (browser) | $0.006/min |
| Latency | < 100ms | ~300ms |
| Offline | ✓ Yes | ✗ No |
| Accuracy | Good | Excellent |
| Languages | ~50 | 99+ |
| Custom vocab | ✗ No | ✓ Yes |
Whisper is planned for Pro tier where transcription accuracy at scale matters. Web Speech API handles real-time practice sessions with zero cost and near-zero latency — the right tradeoff for free-tier users.
What I Built
Designed the Weak Area Memory System — after each session, answers scoring below 6 are flagged, their topics extracted, and stored in the user profile. The next session's Claude prompt is injected with those weak topics so at least 40% of questions target the user's specific gaps, not generic ones
Built a voice interview mode using Web Speech API for browser-based speech-to-text and ElevenLabs for AI voice output — the result is a realistic interview simulation where the AI speaks questions and the user answers by talking, with transcripts scored in real time
Architected a full AWS cloud recording pipeline: MediaRecorder captures audio in-browser, a presigned S3 URL is generated server-side (15-min expiry), the client uploads directly to S3, and CloudFront serves playback globally with per-user scoped access
Implemented the STAR Framework scoring engine using Claude — each answer is evaluated across Situation, Task, Action, and Result (0–2.5 each), with structured JSON output including specific feedback, strength tags, improvement notes, and topic labels for the memory system
Built a Question Prediction Engine: users paste a job description, Claude analyzes the JD language and role signals, and outputs the top 10 most likely interview questions grouped by category — one click starts a session using those exact questions
Designed a GitHub-style Performance Heatmap (52 weeks × 7 days) where each cell's color intensity reflects the user's average score that day, alongside a Topic Coverage Map showing which interview topics are blind spots vs. strengths at a glance
Tech Stack
Frontend
Backend
AI / Voice
Database
Cloud
Platform
From Prototype to Platform
How to Scale This to a Product
IntervAI ships as a feature-complete product. Here's the architecture needed to evolve it from a solo platform into a distributed, multi-tenant SaaS serving 100k+ active users — without changing the interface.
Event-Driven Scoring Pipeline
Apache KafkaInterview sessions emit events (answer.submitted, session.ended) consumed by independent scoring workers and memory-update consumers via Kafka topics. Zero coupling between ingestion and processing — consumer groups absorb traffic spikes with backpressure handling and horizontal fan-out to analytics.
CQRS + Materialized Projections
Read/Write SeparationCommand model: answer submissions hit the PostgreSQL write path through command handlers. Query model: heatmaps, weak-area dashboards, and leaderboards served from Redis-backed materialized projections updated asynchronously. Analytics queries never contend with live interview sessions.
Multi-Tenant B2B Architecture
Schema-Per-TenantEnterprise mode for universities and bootcamps: schema-per-tenant PostgreSQL isolation, Row-Level Security (RLS) enforced at the DB layer, tenant context propagated via signed JWT claims. Enables white-label deployments with full data isolation and zero application-layer config changes per customer.
LLM Gateway + Token Budgeting
AI Cost ControlAll Claude API calls proxied through an internal LLM gateway enforcing per-user token quotas. Prompt caching eliminates redundant context re-tokenization across sessions. Automatic model fallback routing to Haiku on budget exhaustion — zero UX disruption. Per-request cost attribution feeds usage-based Stripe billing.
Circuit Breakers + Bulkhead Isolation
Resilience4JEvery external AI call wrapped in a Resilience4J circuit breaker with half-open state probing. On Claude API degradation, scoring falls back to cached rubric templates. ElevenLabs voice synthesis isolated in a dedicated bulkhead thread pool — a slow TTS request never starves scoring requests of threads.
Autoscaling on Kafka Consumer Lag
Kubernetes HPASpring Boot scoring pods scale on Kafka consumer lag — not CPU. When interview submission spikes hit (exam season, morning rushes), custom HPA metrics trigger scale-out before latency degrades. Zero-downtime rolling deploys with readiness probe gating prevent in-flight sessions from being dropped mid-interview.
Global Voice CDN + Content Caching
CloudFront EdgeElevenLabs TTS responses cached at CloudFront edge nodes keyed by prompt content hash. Identical interview questions served from CDN instead of re-synthesized — cuts TTS API costs ~60% at scale. S3 presigned URLs (15-min expiry) with per-user CloudFront signed cookies scope recording playback access.
Observability with SLI/SLO Contracts
OpenTelemetryDistributed traces propagated via W3C TraceContext across all services. Custom SLOs: scoring_latency_p99 < 2s, session_completion_rate > 85%, weak_area_injection_accuracy > 70%. PostHog funnels track free→paid conversion; Sentry captures AI parse failures with full session context for replay.
Explore the project