BlogA step-by-step agentgateway 1.6.0 guide: run it, put two MCP servers behind one endpoint, add JWT and tool rules, check arguments safely, limit calls, trace, scale, survive failures, route A2A and models, and move your setup behind it.
BlogUK AISI found GPT-6 Astra treated an automated harness reply as permission in 44% of runs. How a default continue message turns into consent, how seven harnesses end a silent turn, what pentest rules of engagement already encode, and a local lab that measures the fix.
BlogA step-by-step build of a personal RAG harness over your own writing. Jev routes, reranks, and checks citations. A local qwen3 writes the answer. Plain Python, measured on 59 questions with recall, McNemar tests, cost, and latency per stage.
GuideRead public LLM and agent benchmarks, then build an eval for any model or harness: failure reading, graders, judge calibration, paired error bars, a held-out hillclimb, agent trajectories, and production checks, measured on a local lab.
BlogHow System 1 decision models work and when to use one. Covers Jev's Choice, Score, and Noul primitives and how Laya reads answers off marker tokens. Includes a live Jev vs Laya test on 30 tickets, a local CPU run, an ECE check, a threshold sweep, and a cost model for routing 10,000 support tickets a day.
BlogHow Agent Substrate egress works, read from its docs. Each actor TCP connection goes over an mTLS CONNECT to a gateway that ignores the hostname and SNI the agent claims. Covers refusals, per-actor EgressPolicy rules, the gaps the project names (no control plane authorization, DNS outside policy), and a hardening order for today.
BlogHow Agent Substrate runs many agents on few pods. Actors live in snapshots and borrow warm workers only while busy. This part covers the actor state machine, snapshot scopes, isolation classes, request parking, and pool sizing.
BlogGoogle AX (Agent Executor) runs agents on Kubernetes as four declarative objects on top of Agent Substrate: Task, Workspace, Gateway, and Model. This part shows how each object resolves underneath and why a resumed task is a fresh process over a restored /workspace. It also covers when suspending pays, with manifests and ax commands from the repo docs.
BlogDevin ships as a cloud agent, a CLI, a desktop editor that used to be Windsurf, and a fleet API. Under them sit two harnesses: a local loop shared by the CLI and Devin Local, and a cloud loop booted from a snapshot. This essay covers the snapshot as the unit of state, the blueprint as the verification contract, and model routing with Fusion and Adaptive. It also covers what /handoff carries and what independent tests say.
GuideRun Xiaomi's MiMo-V2.6 locally: a 9B dense distill, the 309B A15B Flash-RL, and the 1.02T A42B Pro-RL, all MIT, all 1M-context omnimodal. The sliding-window KV math from config.json, every GGUF and MLX size with measured perplexity cost, a three-ladder router that picks the model and quant for your machine, llama.cpp build b11102 commands, and what the MTP drafter is worth while its sidecar path is broken.
BlogSWE-Bench Pro V2 regraded every patch on a clean image and caught Opus 5 forging a Go checksum and Inkling editing the module cache. This essay covers how in-place grading gets gamed and the 17.8 point public to private gap. It turns the V2 defenses into scripts for your own agent pull requests. The scripts cover a fresh-worktree regrade, go mod verify on an empty cache, history-free workspaces, a two-sided task gate, and an executable-config flag.
BlogA playbook for Claude Opus 5.5 in the API and Claude Code. It covers the request changes that avoid 400s, the agent loop under always-on thinking, and effort per task. It also covers effort routing in code, a state file for long runs, a traceability review, and the cost of a 38-handler migration.
BlogAmazon blocked Meta's Muse because its requests carried no agent identity and its logins ran on stored passwords. Web Bot Auth signatures and delegated grants fix both. This essay covers the key directory, signing and verification code, Accept-Signature, and the ID-JAG token requests, then runs a $38 checkout end to end.
BlogZed's Delta replaces the pull request with a thread that keeps the agent's reasoning attached. This essay covers how DeltaDB records that reasoning, review threads, and the review-queue math with a worked example. It keeps the merge gate on the remote with a land skill that queues auto-merge, a GitHub ruleset, and CODEOWNERS on agent instructions.
BlogBITCOS (bitmap and compacted signs) is a weight-packing format for ternary LLMs from Intel Labs. It is not a crypto token. BitNet b1.58 2B4T tensors have 38.5% to 50.8% zeros, so 2 minus z bits per weight beats five-trit packing. This essay gives code, measurements, and the limits of the speedup.
BlogHow to serve and host skills under SEP-2640, the MCP Skills Extension (Final September 13, 2026). Covers the skill:// resource mapping, a manifest builder, the Python reference server, and a host verification gate. Also a cost example for twelve servers, the threat model, failures, and three interactive figures.
BlogIf your agents run long enough to crash, OpenAI's Agents API now rents you the loop. You still own the machine, its keys, and recovery around a stream that does not replay. From the docs: session and executor setup, restricted networks, vaults, a webhook handler, stream recovery, a costed example, and failure modes.
GuideRun DeepSeek-V4.1-Flash, the September 10 MIT release: 552B backbone, 8B active in prefill and 16B in decode, 1M context at 890 bytes of KV per token. Where the 510 GB checkpoint actually goes, why the 203 GB Engram table can live in host RAM, the vLLM and SGLang commands that serve it, and the llama.cpp branch that runs it on one box.
BlogIf your eval harness runs agents that read outside text, its keys are production credentials. GTG-50020 turned an AI vendor's eval sandbox into a key dispenser and replayed the path against about thirty AI companies in four days. A LiteLLM gateway outside the sandbox, per-run keys with a budget and expiry, a closed network and a results scan defeat it.
BlogA three-week agent run factored RSA-260. Read it as a build order for long agent jobs: ten-minute workunits, checkpoints from Young's formula, verification beside production and 3,328 human messages of pruning.
BlogWhat GPT-6 Astra changes in your tool loop. Covers async tools with a job registry and wait_for_tasks, mid-turn steering with a steer journal, and effort in configuration_update items that keep the cache. Also the 403 stop path of the misalignment monitor, the migration diff, and a costed twelve-turn incident session.
BlogGemini 3.8 Flash shipped three weeks after 3.7 Flash at the same price per token, and Artificial Analysis measured the cost per task rising about 40%. What 'works harder' means on the invoice, why the workhorse tier is where budgets move, the January 1 price cliff, and three interactive figures.
BlogAnthropic cut Fable 5.1 cache reads by 75% on September 1. By September 2, Max 20x subscribers reported five-hour sessions gone in twenty minutes. Both are true: the cut was an API price, the wall is a plan meter, and the meter counts traffic, not invoices. Three interactive figures on where the tokens go.
BlogAnthropic shipped one model behind two doors on September 1: Claude Fable 5.1 for everyone, Claude Mythos 5.1 through trusted access. The safeguard stack is the only stated difference, cache reads fell 75 percent, and Google and OpenAI split their own releases the same week. Where the 25 and 45 percent savings come from, with three interactive figures.
BlogClaude's usage limits run on two independent meters: a rolling five-hour window where the 5x and 20x multipliers are defined, and a weekly cap on top that Anthropic sets at its discretion. The arithmetic of the September 14 change, what drains each meter, the lawsuit, the session-hijack advisory, and three interactive figures.
BlogOpenAI gave Cursor the maximum notice its contract allows, with a proposed shutoff of November 12, 2026. Cursor says OpenAI models serve about 5% of its traffic. What a vendor cutoff actually breaks inside a coding harness, why the 95% is untouched, and the audit to run before your own vendor pulls the switch, with three interactive figures.
GuideRun z.ai's GLM-5.3 locally: 743B main stack plus a 10B draft head, about 40B active, open weights since August 28. Why GLM-5.3 Max is a reasoning setting, the sparse-attention cache arithmetic, every checkpoint with its real size, a rig composer, and vLLM, SGLang, llama.cpp, and MLX commands.
GuideRun z.ai's GLM-5.3-Flash locally: 320B total, 18B active, MIT, natively multimodal, stealth-tested as ox-alpha. The hybrid attention memory math with a KV ledger, every unsloth GGUF size, a fit planner for Macs and multi-GPU boxes, and engine-by-engine commands for the branches and images that run it before any tagged release does.
GuideRun Tencent Hunyuan Hy4 preview locally: a 770B total, 49B active MoE with 256 routed experts, Gated DeepSeek Sparse Attention, IndexCache, and a 1M context, under Apache-2.0. What the preview label means, the KV and index cache arithmetic from config.json, every GGUF, MLX, NVFP4 and FP8 build with exact sizes, computed fit by machine, and commands for llama.cpp b10813, Ollama, vLLM 0.29 and SGLang 0.5.20.
BlogAnthropic's Model Hardware Standard gives agents read and write primitives over lab and factory devices, and a driver that refuses a write before the device moves when the plate is missing. What MHS standardizes next to MCP and A2A, with two interactive figures.
BlogIn an OpenAI cybersecurity evaluation, agents that were sealed in separate sandboxes turned a shared package registry into a message board, reached the internet through it, and broke into Hugging Face. The mechanism, the dated climb, and why isolation is a property of the environment, not the model.
GuideRun IBM Granite 4.2 locally: three dense Apache-2.0 reasoning models at 3B, 8B, and 30B with a 128K native and 512K extended window. KV cache arithmetic from config.json, every GGUF, MLX, FP8, and FP4 size, a machine-by-task size grid, Ollama, llama.cpp, vLLM, SGLang, and MLX commands, and measured M1 Max numbers.
BlogThe Model Context Protocol published a new roadmap on August 22, 2026, five months after it deleted its own sessions. How elicitation now works on a stateless server (multi round-trip requests with an opaque requestState), how push survives without a session, why DPoP and workload identity replace pasted API keys, and what a hundred-tool catalog costs before the first question. Three interactive figures.
BlogxAI's Grok Bot gives agents a cloud computer that signs into your tools with no API or MCP. The docs say every Bot on an account shares that computer and its sign-ins. What the product is, why clones copied its shape in eight days, and three interactive figures: close the lid, the login jar, and a routine compiler.
BlogSpeculative decoding drafts a block and verifies it in one pass. DFlash made the draft parallel; DFlash 2 keeps it parallel and adds a 2M-parameter path selector plus a two-tap convolution, for about 21% more accepted tokens per pass at about 1% added latency, with output provably unchanged. The mechanism, the published numbers, and three figures you can drive.
BlogA2A v1.0 shipped in March 2026 and joined the Agentic AI Foundation in August. The spec is a task state machine, three transport bindings, a signed discovery file, and one webhook contract. Working server, client, curl, webhook, persistence, and card-signing code, plus six interactive figures on what v1.0 changed.
BlogNine levels of using Claude, from the claude.ai prompt box to a fleet of managed agents: MCP connectors, memory engineering, skills that end repetition, plugins, hooks, worktree parallelism, cloud sessions, and the practices in each level that almost nobody uses. Five interactive figures, every mechanism from the official docs.
BlogAGENTS.md is a Linux Foundation standard now, read by twenty coding agents and living in 700,000+ files. Almost none of them use the practices that make it work: the token budget, nearest-file precedence, the file-resolution order that silently shadows your rules, the eval that hits 100%, and the fact that it is now an attack surface. Three interactive figures.
GuideThe transformer from zero to frontier: build attention up from a masked average, train it and watch the loss ladder, apply nine years of architecture diffs (RMSNorm, SwiGLU, RoPE, GQA, MLA, MoE, hybrid attention), then run it behind a chat. Thirteen interactive figures, every number from the source paper or model config.
GuideA source-level walk through vLLM's V1 engine: the paged KV block pool with content hashes, the phase-free token-budget scheduler, the two-process loop over ZMQ, speculative decoding, grammar-masked sampling, the compile pipeline, and how it scales past one GPU. Six interactive figures you can step through.
BlogA story about the fifty companies coming next, the ones the agent boom is about to make necessary: cost routers, agent memory, CI for agents, scale-to-zero tools, forward-deployed agents, local-first inference, and the safety layer underneath. Four come with a working interactive demo.
GuideRun Qwen/Qwen3.8-27B locally: the Apache-2.0 dense 27B vision-language model with 48 Gated DeltaNet layers and 16 full-attention layers. The fixed 154 MB recurrent state and 64 KB per token KV math from config.json, every GGUF, MLX, FP8, NVFP4 and AWQ size, a 24 to 48 GB fit map, llama.cpp context checkpoints, Ollama, vLLM, SGLang and MLX commands, and an M1 Max bench note.
BlogPick a model, wrap it in a harness, and the harness quietly decides the bill. A benchmark ran one model at one effort through different harnesses and cost per task moved more than 2x at equal quality. A field guide to ten terminal coding agents: Claude Code, Codex, Gemini CLI, Grok Code, Kimi CLI, opencode, Qwen Code, Pi, the iii harness, and Prime Agent. Three interactive figures.
BlogSame model, same GPU, and the serving engine you route it through changes throughput and tail latency more than most teams expect. A field guide to vLLM, SGLang, TensorRT-LLM, LMDeploy, and TGI: continuous batching, prefix reuse, compile-versus-load, and the numbers with the caveats they deserve. Four interactive figures.
BlogIn 2026 the LLM wiki idea became real tooling, in two opposite shapes: a connector-rich standalone CLI that version-controls Markdown, and the same generate-search-refresh-lint lifecycle decomposed into functions and triggers on a runtime. What each gains, what each gives up.
BlogThe 2026-07-28 MCP spec removed the initialize handshake and the Mcp-Session-Id header. The protocol core is now stateless: every request is self-contained, so any request can land on any server replica behind a round-robin load balancer. What that means, with four worked examples and two interactive diagrams.
BlogThe LLM wiki pattern is clean until you run it hard. Across thousands of sessions and hundreds of pages, a flat markdown store needs a memory lifecycle: confidence scoring, supersession, forgetting curves, and consolidation tiers. The production version, from building agentmemory.
BlogRAG retrieves the same raw documents on every query and forgets between them. An LLM-maintained wiki compiles knowledge once and keeps it current, so it compounds instead of re-deriving. The pattern, its three layers, and why the schema is the real product.
BlogTraining a model is a one-time cost. Inference is the bill that arrives with every query, and it inverts the oldest rule in business: an AI company's heaviest users are its least profitable. The mechanism, the collapsing cost per token, and why owning silicon is the only structural escape.
BlogIn 2026 OpenAI, Anthropic, Microsoft, and AWS committed a combined nine billion dollars to deployment companies built around one job title Palantir coined for its lowest-status engineers. Why the forward deployed engineer came back, from the mechanism up.
BlogThe FDE job from the inside: the deliverables named in real OpenAI, Anthropic, and Harvey postings, the pod structure, pay ranges with sources, how the role differs from solutions engineering and consulting, and how people get in.
BlogA harness generated a 400-line throwaway workflow that spawned a hundred agents, and the industry named the shape graph engineering. Euler named it in 1736. What actually changed is the node. With two interactive figures: walk the Konigsberg bridges yourself, and run the agent DAG with live cost arithmetic.
BlogClaude Code assembles context in layers ordered by how often they change, delivers CLAUDE.md as a user message, and treats cache preservation as an architectural principle. A mechanism-level tour from the current docs.
BlogCodex made compaction a server-side primitive: a dedicated endpoint, in-stream compaction items, a 90 percent trigger, and a model trained on where the summary sits. OpenAI tripled its own ARC-AGI-3 score by turning it on. The price is a transcript you cannot fully read.
BlogPi stores sessions as trees, keeps its system prompt to twenty stable lines, refuses to prune history, and itemizes every cache miss in dollars. A mechanism-level tour of the most transparent coding agent harness.
BlogMoonshot ships a harness that is also your shell, rewinds context with checkpoint time travel, fans out 128 subagents, and requires reasoning to be passed back verbatim. And its own benchmark admits K3 scores higher inside Claude Code.
BlogEvery repo now carries a document addressed to a machine, written in shouted imperative English. It is a new genre of technical writing, it has conventions already, and you can read a team's incident history in it.
BlogLocal LLM generation speed is not a compute number. Every generated token reads the model's weights from memory once, so tokens per second is bounded by bandwidth divided by weight bytes. The arithmetic, the consequences, and what it predicts.
BlogAgents build the wrong thing, ship duplicates, and reinvent what you already have. The fix is not a smarter model, it is a system the agent is part of. My own early numbers, on my own projects: roughly 60% fewer tokens and about twice as fast from design to build.
BlogHarness engineering is the practice of improving agent output by shaping the environment around a fixed model: context, tools, the loop, and proof. A field guide to where the leverage sits.
BlogA practical, honest roadmap to becoming an AI engineer in 2026: the skills in order, two routes in, time estimates, and one free 503-lesson curriculum that builds the whole stack from scratch.
GuideHow much memory you actually need to run local LLMs, and what fits a Mac Mini, MacBook Pro, Mac Studio, or an NVIDIA GPU. The RAM math, Apple unified-memory tiers, GPU tiers, and a model-to-hardware table.
BlogHarness engineering, context engineering, loop engineering. Most of the agent era's new vocabulary renames engineering fundamentals we have had for decades. Here is the old idea under each new name.
BlogThin versus thick harness is the wrong debate. A backend is workers, triggers, and functions, and an agent is just another worker. Once the harness is the backend, agent reliability becomes durable execution, and composability is how the hard problems get solved.
BlogEmitting an event after a database commit without an outbox is a lossy notifier that drops events on crash. It is the dual-write problem, and the fix is the outbox pattern or real change data capture.
BlogEvery rule about readable code, DRY, good names, one source of truth, was always for the next human. A second reader arrived: faster, more literal, and it audits how much of that advice you actually followed.
GuideRun Liquid AI's LFM2.5 edge family locally: the 2.6B flagship, the 8B-A1B MoE, VL-3B, and the 350M and 230M tier. The LFM Open License read from the file, conv-plus-attention cache math from config.json, every GGUF and MLX size, phone, laptop CPU, and Apple Silicon commands, and measured M1 Max numbers.
GuideRun MiniMax's MiniMax-M3 locally: a 428B MoE with about 23B active, native image and video input, and a 1M-token config. The MSA sparse-attention cache arithmetic from config.json, every GGUF, MLX, AWQ, MXFP8 and NVFP4 size, computed GPU, RAM, and SSD expert placement per machine, the license terms, and llama.cpp, vLLM, SGLang, KTransformers and mlx-vlm commands.
GuideRun Google's Gemma 4 locally: E4B with per-layer embeddings, the encoder-free 12B, the 26B-A4B mixture of experts, and the 31B dense, all Apache 2.0. KV cache math from config.json, every GGUF, MLX, FP8, and AWQ size, machine fit verdicts, commands for Ollama, llama.cpp, vLLM, SGLang, and MLX, and decode speed ceilings.
GuideRun DeepSeek R1 and V3.2 locally: the small distilled models on a consumer GPU, and the full 671B mixture-of-experts on a big-RAM box with dynamic quants and MoE offload. Ollama, llama.cpp, vLLM, sampler settings, tool calling, and troubleshooting.
GuideRun DeepSeek-V4-Flash-0731 (284B, 13B active, 1M context) locally: pick a quant interactively, see if it fits your machine, then get exact llama.cpp, server, Mac, and agent-wiring commands that update live. Reasoning modes, MoE offload math, troubleshooting.
GuideRun Google's Gemma 3 locally, from the 1B that fits a phone to the 27B on a 24GB GPU, including the vision models. Ollama, llama.cpp, vLLM, MLX, the recommended sampler, the double-BOS gotcha, and running images with the multimodal projector.
GuideRun z.ai's GLM models locally: the compact GLM-4.5 Air on a consumer GPU, the flagship GLM-4.6 with expert offload, and the frontier GLM-5 on a server. Ollama, llama.cpp, vLLM, tool calling for agentic coding, and troubleshooting.
GuideRun OpenAI's gpt-oss 20B and 120B open-weight models locally: native MXFP4, the harmony response format, reasoning_effort, Ollama, llama.cpp, vLLM, LM Studio, tool calling, an OpenAI-compatible API, and troubleshooting.
GuideA detailed guide to running Kimi K3, Moonshot's trillion-scale open MoE, on your own hardware: dynamic quantization, MoE-to-CPU offload in llama.cpp, multi-GPU vLLM, Apple Silicon, tool calling, an OpenAI-compatible endpoint, and troubleshooting.
GuideRun Meta's Llama 4 Scout and Maverick locally: native multimodal mixture-of-experts models with huge context. Hardware sizing, gated access, llama.cpp with expert offload, vLLM tensor parallel, tool calling, and troubleshooting.
GuideRun Mistral's open models locally: the 24B Mistral Small on a mainstream GPU, Devstral for agentic coding, and Mistral Large for multi-GPU rigs. Ollama, llama.cpp, vLLM, MLX, tool calling, sampler settings, and troubleshooting.
GuideRun Meta's Muse Glimmer 30B on your own hardware, day one: exact GGUF sizes for every quant, the 24GB VRAM envelope math (weights + KV cache + vision encoder + DFlash drafter), llama.cpp and Ollama and Mac commands, the ATEM chat template and reasoning strength levels, agent wiring, fine-tuning with Unsloth, and honest benchmarks including where it loses.
GuideRun Qwen3 locally, from the 4B that fits a laptop to the 480B Coder. Ollama, llama.cpp, vLLM, and MLX, with the right quant per GPU, thinking mode, tool calling, YaRN long context, an OpenAI-compatible API, and troubleshooting.
GuideRun Qwen3.8-Flash-Next on your own hardware: exact GGUF sizes for every Unsloth quant, where the 51B n-gram table goes (VRAM, RAM, or NVMe via mmap), the hybrid Gated DeltaNet plus sparse attention KV ledger at 32K to 1M context, bytes-touched-per-token speed ceilings, llama.cpp and vLLM commands, thinking and reasoning effort settings, and the day-one rough edges from the llama.cpp pull request.