Blog
Recent Articles
Discover the latest news, tips and system design from Codemia.
You will learn about web infrastructure, system designs and devops APIs best practices.
Decoding one token reads every weight out of memory, which is why precision is a latency knob before it is a memory knob. A full walkthrough of quantization and distillation, from what a number actually is, through outliers, SmoothQuant, GPTQ, AWQ and KV cache quantization, to soft labels, on-policy distillation, and the order you should apply them in.
By Codemia ⢠Sep 25, 2026
28 min read
Coding agents now read a repository, plan a change, edit several files, run the tests, and iterate until they pass. A full walkthrough of the design, from requirements to sandboxing to evaluation, and what an interviewer is listening for at each step.
By Codemia ⢠Aug 31, 2026
18 min read
Why QUICâs user-space transport lets us âkillâ the old app-level event loop: a tour from Tahoe/Reno to Cubic/BBR, then into QUICâs pacing, loss recovery, and stream-level multiplexing that trim tail latency and jitter at scale.
By Codemia ⢠Oct 9, 2025
10 min read
From MySQLâs retired query cache to modern Redis/Memcached, this deep dive compares DB-level vs app-level query caching with real-world hit/stale rates and practical guidance.
By Codemia ⢠Oct 4, 2025
5 min read
Loading...