VoiceDiary
Classroom Lecture AI & Diarization Engine
GitHub Repo
Real-Time Bilingual Lecture Diarization (Urdu + English)

Automated Classroom Lecture Notes
with AI Speaker Identification

VoiceDiary leverages Faster-Whisper, SpeechBrain ECAPA-TDNN, and Silero VAD to capture, transcribe, and color-code multi-speaker lectures into structured notes with 0% latency.

Live Classroom Lecture Simulation

Urdu + English Code-Switching with Real-Time Speaker Diarization

ENGINE ACTIVE
Professor (Speaker 1) [00:04]

Today we will discuss gradient descent and backpropagation in deep neural networks. Yeh concept bohot critical hai for multi-layer perceptron training.

Student (Speaker 2) [00:21]

Sir, what is the mathematical difference between Stochastic Gradient Descent and Adam Optimizer?

Professor (Speaker 1) [00:32]

Good question. Adam uses adaptive moment estimation, meaning it tracks both first and second moments of the gradients to adjust learning rates dynamically.

Hardware-Adaptive AI Engine

Dynamically binds NVIDIA Tensor Cores (float16) or scales Multi-Core AVX2 INT8 vectorization for 0.2s live response on any PC.

50-Voiceprint Diarization

ECAPA-TDNN cosine matrix centroid matching identifies repeating speakers in <0.1ms with clean visual color tags.

1-Click Multi-Format Export

Instantly export classroom lectures to Markdown notes (.md), plain text transcripts (.txt), or timed subtitles (.srt).