Beyond the curriculum

Blogs & Guides

Technical deep dives by Rohit Ghumare on AI agents, LLM inference, RAG, and developer tools.

Examine how systems work, where they fail, and which tradeoffs matter. Build the foundations in the course:

Explore the latest writing

Blog

Laya and Jev: System 1 Models Answer With a Type, Not Text

How System 1 decision models work and when to use one. Covers Jev's Choice, Score, and Noul primitives and how Laya reads answers off marker tokens. Includes a live Jev vs Laya test on 30 tickets, a local CPU run, an ECE check, a threshold sweep, and a cost model for routing 10,000 support tickets a day.

Blog

Agent Substrate Egress and Its Gaps: Nothing the Agent Says Is Trusted

How Agent Substrate egress works, read from its docs. Each actor TCP connection goes over an mTLS CONNECT to a gateway that ignores the hostname and SNI the agent claims. Covers refusals, per-actor EgressPolicy rules, the gaps the project names (no control plane authorization, DNS outside policy), and a hardening order for today.

Blog

Google AX (Agent Executor): Task, Workspace, Gateway, Model

Google AX (Agent Executor) runs agents on Kubernetes as four declarative objects on top of Agent Substrate: Task, Workspace, Gateway, and Model. This part shows how each object resolves underneath and why a resumed task is a fresh process over a restored /workspace. It also covers when suspending pays, with manifests and ax commands from the repo docs.

Blog

Inside the Devin Harness: Two Loops and a Handoff

Devin ships as a cloud agent, a CLI, a desktop editor that used to be Windsurf, and a fleet API. Under them sit two harnesses: a local loop shared by the CLI and Devin Local, and a cloud loop booted from a snapshot. This essay covers the snapshot as the unit of state, the blueprint as the verification contract, and model routing with Fusion and Adaptive. It also covers what /handoff carries and what independent tests say.

Guide

How to Run MiMo-V2.6 Locally

Run Xiaomi's MiMo-V2.6 locally: a 9B dense distill, the 309B A15B Flash-RL, and the 1.02T A42B Pro-RL, all MIT, all 1M-context omnimodal. The sliding-window KV math from config.json, every GGUF and MLX size with measured perplexity cost, a three-ladder router that picks the model and quant for your machine, llama.cpp build b11102 commands, and what the MTP drafter is worth while its sidecar path is broken.

Blog

SWE-Bench Pro V2 Caught Two Models Cheating. The Fix Was Moving the Grader.

SWE-Bench Pro V2 regraded every patch on a clean image and caught Opus 5 forging a Go checksum and Inkling editing the module cache. This essay covers how in-place grading gets gamed and the 17.8 point public to private gap. It turns the V2 defenses into scripts for your own agent pull requests. The scripts cover a fresh-worktree regrade, go mod verify on an empty cache, history-free workspaces, a two-sided task gate, and an executable-config flag.

Blog

Opus 5.5 in Practice: Effort Is the Only Dial Left

A playbook for Claude Opus 5.5 in the API and Claude Code. It covers the request changes that avoid 400s, the agent loop under always-on thinking, and effort per task. It also covers effort routing in code, a state file for long runs, a traceability review, and the cost of a 38-handler migration.

Blog

Amazon Blocked Muse. The Request Had No Name on It.

Amazon blocked Meta's Muse because its requests carried no agent identity and its logins ran on stored passwords. Web Bot Auth signatures and delegated grants fix both. This essay covers the key directory, signing and verification code, Accept-Signature, and the ID-JAG token requests, then runs a $38 checkout end to end.

Blog

Zed Turned Off Pull Requests. The Merge Gate Has to Go Somewhere.

Zed's Delta replaces the pull request with a thread that keeps the agent's reasoning attached. This essay covers how DeltaDB records that reasoning, review threads, and the review-queue math with a worked example. It keeps the merge gate on the remote with a land skill that queues auto-merge, a GitHub ruleset, and CODEOWNERS on agent instructions.

Blog

BITCOS: Intel's Ternary LLM Weight Format Goes Below 1.58 Bits

BITCOS (bitmap and compacted signs) is a weight-packing format for ternary LLMs from Intel Labs. It is not a crypto token. BitNet b1.58 2B4T tensors have 38.5% to 50.8% zeros, so 2 minus z bits per weight beats five-trit packing. This essay gives code, measurements, and the limits of the speedup.

Blog

MCP Servers Can Now Ship Their Own Manual

How to serve and host skills under SEP-2640, the MCP Skills Extension (Final September 13, 2026). Covers the skill:// resource mapping, a manifest builder, the Python reference server, and a host verification gate. Also a cost example for twelve servers, the threat model, failures, and three interactive figures.

Blog

The Harness Is an API Now. Here Is What You Still Own.

If your agents run long enough to crash, OpenAI's Agents API now rents you the loop. You still own the machine, its keys, and recovery around a stream that does not replay. From the docs: session and executor setup, restricted networks, vaults, a webhook handler, stream recovery, a costed example, and failure modes.

Guide

How to Run DeepSeek-V4.1-Flash: 510 GB, and 203 GB of It Is a Lookup Table

Run DeepSeek-V4.1-Flash, the September 10 MIT release: 552B backbone, 8B active in prefill and 16B in decode, 1M context at 890 bytes of KV per token. Where the 510 GB checkpoint actually goes, why the 203 GB Engram table can live in host RAM, the vLLM and SGLang commands that serve it, and the llama.cpp branch that runs it on one box.

Blog

The Eval Sandbox Held the Keys

If your eval harness runs agents that read outside text, its keys are production credentials. GTG-50020 turned an AI vendor's eval sandbox into a key dispenser and replayed the path against about thirty AI companies in four days. A LiteLLM gateway outside the sandbox, per-run keys with a budget and expiry, a closed network and a results scan defeat it.

Blog

GPT-6 Astra Changed the Shape of a Turn

What GPT-6 Astra changes in your tool loop. Covers async tools with a job registry and wait_for_tasks, mid-turn steering with a steer journal, and effort in configuration_update items that keep the cache. Also the 403 stop path of the misalignment monitor, the migration diff, and a costed twelve-turn incident session.

Blog

Fable 5.1 Burned the Meter in Twenty Minutes

Anthropic cut Fable 5.1 cache reads by 75% on September 1. By September 2, Max 20x subscribers reported five-hour sessions gone in twenty minutes. Both are true: the cut was an API price, the wall is a plan meter, and the meter counts traffic, not invoices. Three interactive figures on where the tokens go.

Blog

Fable 5.1 and Mythos 5.1: One Model, Two Doors

Anthropic shipped one model behind two doors on September 1: Claude Fable 5.1 for everyone, Claude Mythos 5.1 through trusted access. The safeguard stack is the only stated difference, cache reads fell 75 percent, and Google and OpenAI split their own releases the same week. Where the 25 and 45 percent savings come from, with three interactive figures.

Blog

Claude Has Two Meters. The 20x Is Only on One.

Claude's usage limits run on two independent meters: a rolling five-hour window where the 5x and 20x multipliers are defined, and a weekly cap on top that Anthropic sets at its discretion. The arithmetic of the September 14 change, what drains each meter, the lawsuit, the session-hijack advisory, and three interactive figures.

Blog

OpenAI Cut Cursor Off. Five Percent Is the Interesting Number.

OpenAI gave Cursor the maximum notice its contract allows, with a proposed shutoff of November 12, 2026. Cursor says OpenAI models serve about 5% of its traffic. What a vendor cutoff actually breaks inside a coding harness, why the 95% is untouched, and the audit to run before your own vendor pulls the switch, with three interactive figures.

Guide

How to Run GLM-5.3-Flash Locally

Run z.ai's GLM-5.3-Flash locally: 320B total, 18B active, MIT, natively multimodal, stealth-tested as ox-alpha. The hybrid attention memory math with a KV ledger, every unsloth GGUF size, a fit planner for Macs and multi-GPU boxes, and engine-by-engine commands for the branches and images that run it before any tagged release does.

Guide

How to Run Hy4 Preview Locally

Run Tencent Hunyuan Hy4 preview locally: a 770B total, 49B active MoE with 256 routed experts, Gated DeepSeek Sparse Attention, IndexCache, and a 1M context, under Apache-2.0. What the preview label means, the KV and index cache arithmetic from config.json, every GGUF, MLX, NVFP4 and FP8 build with exact sizes, computed fit by machine, and commands for llama.cpp b10813, Ollama, vLLM 0.29 and SGLang 0.5.20.

Guide

How to Run Granite 4.2 Locally

Run IBM Granite 4.2 locally: three dense Apache-2.0 reasoning models at 3B, 8B, and 30B with a 128K native and 512K extended window. KV cache arithmetic from config.json, every GGUF, MLX, FP8, and FP4 size, a machine-by-task size grid, Ollama, llama.cpp, vLLM, SGLang, and MLX commands, and measured M1 Max numbers.

Blog

MCP's New Roadmap: Events Without Sessions, Identity Without API Keys

The Model Context Protocol published a new roadmap on August 22, 2026, five months after it deleted its own sessions. How elicitation now works on a stateless server (multi round-trip requests with an opaque requestState), how push survives without a session, why DPoP and workload identity replace pasted API keys, and what a hundred-tool catalog costs before the first question. Three interactive figures.

Blog

Grok Bot: The Agent Got a Computer. Your Bots Share It.

xAI's Grok Bot gives agents a cloud computer that signs into your tools with no API or MCP. The docs say every Bot on an account shares that computer and its sign-ins. What the product is, why clones copied its shape in eight days, and three interactive figures: close the lid, the login jar, and a routine compiler.

Blog

DFlash 2: The Drafter Stopped Guessing Alone

Speculative decoding drafts a block and verifies it in one pass. DFlash made the draft parallel; DFlash 2 keeps it parallel and adds a 2M-parameter path selector plus a two-tap convolution, for about 21% more accepted tokens per pass at about 1% added latency, with output provably unchanged. The mechanism, the published numbers, and three figures you can drive.

Blog

All About the A2A Protocol, Now Part of the Agentic AI Foundation

A2A v1.0 shipped in March 2026 and joined the Agentic AI Foundation in August. The spec is a task state machine, three transport bindings, a signed discovery file, and one webhook contract. Working server, client, curl, webhook, persistence, and card-signing code, plus six interactive figures on what v1.0 changed.

Blog

Claude Has Levels. Most People Stop at Two.

Nine levels of using Claude, from the claude.ai prompt box to a fleet of managed agents: MCP connectors, memory engineering, skills that end repetition, plugins, hooks, worktree parallelism, cloud sessions, and the practices in each level that almost nobody uses. Five interactive figures, every mechanism from the official docs.

Blog

The AGENTS.md Practices Nobody Uses

AGENTS.md is a Linux Foundation standard now, read by twenty coding agents and living in 700,000+ files. Almost none of them use the practices that make it work: the token budget, nearest-file precedence, the file-resolution order that silently shadows your rules, the eval that hits 100%, and the fact that it is now an attack surface. Three interactive figures.

Guide

Inside the Transformer: Build It, Mutate It, Run It

The transformer from zero to frontier: build attention up from a masked average, train it and watch the loss ladder, apply nine years of architecture diffs (RMSNorm, SwiGLU, RoPE, GQA, MLA, MoE, hybrid attention), then run it behind a chat. Thirteen interactive figures, every number from the source paper or model config.

Guide

Inside the vLLM Engine: How the Default LLM Serving Engine Works

A source-level walk through vLLM's V1 engine: the paged KV block pool with content hashes, the phase-free token-budget scheduler, the two-process loop over ZMQ, speculative decoding, grammar-masked sampling, the compile pipeline, and how it scales past one GPU. Six interactive figures you can step through.

Guide

How to Run Qwen3.8-27B Locally

Run Qwen/Qwen3.8-27B locally: the Apache-2.0 dense 27B vision-language model with 48 Gated DeltaNet layers and 16 full-attention layers. The fixed 154 MB recurrent state and 64 KB per token KV math from config.json, every GGUF, MLX, FP8, NVFP4 and AWQ size, a 24 to 48 GB fit map, llama.cpp context checkpoints, Ollama, vLLM, SGLang and MLX commands, and an M1 Max bench note.

Blog

The Harness, Not the Model: Ten Coding Agents Compared

Pick a model, wrap it in a harness, and the harness quietly decides the bill. A benchmark ran one model at one effort through different harnesses and cost per task moved more than 2x at equal quality. A field guide to ten terminal coding agents: Claude Code, Codex, Gemini CLI, Grok Code, Kimi CLI, opencode, Qwen Code, Pi, the iii harness, and Prime Agent. Three interactive figures.

Blog

The Inference Engine Underneath: vLLM, SGLang, and the Serving Stack

Same model, same GPU, and the serving engine you route it through changes throughput and tail latency more than most teams expect. A field guide to vLLM, SGLang, TensorRT-LLM, LMDeploy, and TGI: continuous batching, prefix reuse, compile-versus-load, and the numbers with the caveats they deserve. Four interactive figures.

Blog

OpenWiki: The LLM Wiki Pattern as Running Code

In 2026 the LLM wiki idea became real tooling, in two opposite shapes: a connector-rich standalone CLI that version-controls Markdown, and the same generate-search-refresh-lint lifecycle decomposed into functions and triggers on a runtime. What each gains, what each gives up.

Blog

Stateless MCP: The Protocol Deleted Its Own Handshake

The 2026-07-28 MCP spec removed the initialize handshake and the Mcp-Session-Id header. The protocol core is now stateless: every request is self-contained, so any request can land on any server replica behind a round-robin load balancer. What that means, with four worked examples and two interactive diagrams.

Blog

LLM Wiki v2: What Breaks at Scale

The LLM wiki pattern is clean until you run it hard. Across thousands of sessions and hundreds of pages, a flat markdown store needs a memory lifecycle: confidence scoring, supersession, forgetting curves, and consolidation tiers. The production version, from building agentmemory.

Blog

The LLM Wiki: Stop Retrieving, Start Compounding

RAG retrieves the same raw documents on every query and forgets between them. An LLM-maintained wiki compiles knowledge once and keeps it current, so it compounds instead of re-deriving. The pattern, its three layers, and why the schema is the real product.

Blog

Graph Engineering Is 290 Years Old

A harness generated a 400-line throwaway workflow that spawned a hundred agents, and the industry named the shape graph engineering. Euler named it in 1736. What actually changed is the node. With two interactive figures: walk the Konigsberg bridges yourself, and run the agent DAG with live cost arithmetic.

Blog

The AGENTS.md Genre

Every repo now carries a document addressed to a machine, written in shouted imperative English. It is a new genre of technical writing, it has conventions already, and you can read a team's incident history in it.

Blog

Tokens per Second Is a Memory Bandwidth Number

Local LLM generation speed is not a compute number. Every generated token reads the model's weights from memory once, so tokens per second is bounded by bandwidth divided by weight bytes. The arithmetic, the consequences, and what it predicts.

Blog

Your Codebase Has a Second Reader Now

Every rule about readable code, DRY, good names, one source of truth, was always for the next human. A second reader arrived: faster, more literal, and it audits how much of that advice you actually followed.

Guide

How to Run LFM2.5 Locally

Run Liquid AI's LFM2.5 edge family locally: the 2.6B flagship, the 8B-A1B MoE, VL-3B, and the 350M and 230M tier. The LFM Open License read from the file, conv-plus-attention cache math from config.json, every GGUF and MLX size, phone, laptop CPU, and Apple Silicon commands, and measured M1 Max numbers.

Guide

How to Run MiniMax-M3 Locally

Run MiniMax's MiniMax-M3 locally: a 428B MoE with about 23B active, native image and video input, and a 1M-token config. The MSA sparse-attention cache arithmetic from config.json, every GGUF, MLX, AWQ, MXFP8 and NVFP4 size, computed GPU, RAM, and SSD expert placement per machine, the license terms, and llama.cpp, vLLM, SGLang, KTransformers and mlx-vlm commands.

Guide

How to Run Gemma 4 Locally

Run Google's Gemma 4 locally: E4B with per-layer embeddings, the encoder-free 12B, the 26B-A4B mixture of experts, and the 31B dense, all Apache 2.0. KV cache math from config.json, every GGUF, MLX, FP8, and AWQ size, machine fit verdicts, commands for Ollama, llama.cpp, vLLM, SGLang, and MLX, and decode speed ceilings.

Guide

How to Run Muse Glimmer Locally: Meta's 30B Local Agent Model

Run Meta's Muse Glimmer 30B on your own hardware, day one: exact GGUF sizes for every quant, the 24GB VRAM envelope math (weights + KV cache + vision encoder + DFlash drafter), llama.cpp and Ollama and Mac commands, the ATEM chat template and reasoning strength levels, agent wiring, fine-tuning with Unsloth, and honest benchmarks including where it loses.

Guide

How to Run Qwen3.8-Flash-Next Locally

Run Qwen3.8-Flash-Next on your own hardware: exact GGUF sizes for every Unsloth quant, where the 51B n-gram table goes (VRAM, RAM, or NVMe via mmap), the hybrid Gated DeltaNet plus sparse attention KV ledger at 32K to 1M context, bytes-touched-per-token speed ceilings, llama.cpp and vLLM commands, thinking and reasoning effort settings, and the day-one rough edges from the llama.cpp pull request.