Applied AI, ML systems, and research notes by Juntak Noh.
This site contains:
- Paper reviews (assumptions → failure modes → what I’d change)
- Applied ML system design notes
- Reproducible experiments
Applied AI, ML systems, and research notes by Juntak Noh.
This site contains:
This is Part 7, the last of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → Metrics → Simplification. Part 5 showed how the dialog layer’s rules piled up, and Part 6 why the evaluation kept rewarding them. This part is about taking them out. ...
This is Part 6 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → Metrics → Simplification. Part 5 covered how the dialog layer’s rules piled up one failing QA row at a time. This part is about the evaluation that kept rewarding them. ...
This is Part 5 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, Part 2 where rules beat models, Part 3 latency, and Part 4 failures that never threw. This part is about the dialog layer, and how its rules multiplied one failing QA row at a time. ...
This is Part 4 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, Part 2 where rules beat models, and Part 3 latency. This part is about the failures that never threw an exception. When a web app breaks, someone at least sees a 500 page. When a voice agent breaks, the caller hears nothing. They say “여보세요?” into the silence, wait a few seconds, and hang up. No stack trace reaches them, and often none reaches you either. ...
This is Part 3 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, and Part 2 where rules beat models. This part is about latency: where the seconds went, and why the fixes that mattered were almost never “use a smaller model.” ...
This is Part 2 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me. This part is about the text around the LLM: numbers going in and out, names that STT gets almost right, and where an LLM call is worth its latency. ...
This is Part 1 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. The Whisper fine-tuning series covered how I trained the model; this one starts with what happened when it met real calls. By the end of last year my Whisper fine-tuning curve looked about as good as curves get. On the held-out split of a public Korean 8 kHz telephone corpus (AI Hub’s low-quality telephone-network speech data), every round beat the last: 9.0% CER, then 6.3%, then 5.5%, then 3.8% for a fine-tuned large-v3-turbo. Off-the-shelf large-v3-turbo sat at 6.5%. I had set myself a target of under 4%, and 3.8% cleared it. ...
Everyone can recite “high bias is underfitting, high variance is overfitting.” Far fewer can look at a training run and say which one they have and what to do about it. This post does both: first the bias–variance decomposition tightly enough to be useful, then a practical playbook — learning curves, the train/validation gap, cross-validation, and the traps — for diagnosing overfitting on a real model. It’s the diagnostic companion to the L1/L2 and dropout posts, which cover the fixes. ...
Dropout is two ideas bolted together: randomly switch off units during training, then quietly turn them all back on for inference. The interesting parts are (1) why switching units off randomly improves generalization at all, and (2) the fact that “turn them all back on” is not free — the activations come out at the wrong scale unless you correct for it. Getting the scaling wrong is one of the most common deep-learning bugs, so this post works through both, at the same engineering depth as the L1/L2 post. ...
“L1 gives you sparse weights, L2 gives you small weights” What Regularization Actually Is A model with enough capacity will happily drive its training loss to zero by memorizing the data — including the noise. That is overfitting: great training numbers, bad predictions on anything new. In a linear model it shows up as coefficients that blow up to large, opposing values. Why does fitting noise require large coefficients? Because of the amplification in the OLS solution \(\hat{\mathbf{w}} = (X^\top X)^{-1}X^\top y\). If \(X^\top X\) has a small eigenvalue (a direction of almost no variance in the data — e.g. two nearly identical features), its inverse has a large eigenvalue, and noise in \(y\) projected onto that direction gets amplified into a large coefficient. Concretely, if \(x_1 \approx x_2\), explaining \(y\) needs only \(w_1=1,\,w_2=0\) — but squeezing out the last bit of noise-fit might use \(w_1=1000,\,w_2=-999\). Since \(x_1 - x_2\) is a near-zero direction, you need enormous coefficients to produce the tiny output that matches the noise. Large \(|\mathbf{w}|\) is the fingerprint of a function contorting itself to fit noise. ...