{"items":[{"path":"certifications/claude/lessons/04-context-knowledge-memory-and-caching","kind":"lesson","title":"Put Each Fact in the Right Kind of Context","description":"Put Each Fact in the Right Kind of Context: Context is temporary attention. Knowledge is maintained evidence. Memory is continuity. Caching is reuse. Mixing…","url":"https://aiengineeringfromscratch.com/lesson?path=certifications%2Fclaude%2Flessons%2F04-context-knowledge-memory-and-caching","sourceUrl":"https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/certifications/claude/lessons/04-context-knowledge-memory-and-caching/docs/en.md"},{"path":"phases/05-nlp-foundations-to-advanced/08-cnns-rnns-for-text","kind":"lesson","title":"CNNs and RNNs for Text","description":"CNNs and RNNs for Text: Convolutions learn n-grams. Recurrences remember. Both are superseded by attention. Both still matter on constrained hardware.","url":"https://aiengineeringfromscratch.com/lesson?path=phases%2F05-nlp-foundations-to-advanced%2F08-cnns-rnns-for-text","sourceUrl":"https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/05-nlp-foundations-to-advanced/08-cnns-rnns-for-text/docs/en.md"},{"path":"phases/05-nlp-foundations-to-advanced/09-sequence-to-sequence","kind":"lesson","title":"Sequence-to-Sequence Models","description":"Sequence-to-Sequence Models: Two RNNs pretending to be a translator. The bottleneck they hit is the reason attention exists. Classification maps a…","url":"https://aiengineeringfromscratch.com/lesson?path=phases%2F05-nlp-foundations-to-advanced%2F09-sequence-to-sequence","sourceUrl":"https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/05-nlp-foundations-to-advanced/09-sequence-to-sequence/docs/en.md"},{"path":"phases/05-nlp-foundations-to-advanced/10-attention-mechanism","kind":"lesson","title":"Attention Mechanism — The Breakthrough","description":"Attention Mechanism — The Breakthrough: The decoder stops squinting at a compressed summary and starts looking at the whole source. Everything after this is…","url":"https://aiengineeringfromscratch.com/lesson?path=phases%2F05-nlp-foundations-to-advanced%2F10-attention-mechanism","sourceUrl":"https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/docs/en.md"},{"path":"phases/06-speech-and-audio/04-speech-recognition-asr","kind":"lesson","title":"Speech Recognition (ASR) — CTC, RNN-T, Attention","description":"Speech Recognition (ASR) — CTC, RNN-T, Attention: Speech recognition is audio classification at every timestep, glued together by a sequence model that knows…","url":"https://aiengineeringfromscratch.com/lesson?path=phases%2F06-speech-and-audio%2F04-speech-recognition-asr","sourceUrl":"https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/06-speech-and-audio/04-speech-recognition-asr/docs/en.md"},{"path":"phases/07-transformers-deep-dive/02-self-attention-from-scratch","kind":"lesson","title":"Self-Attention from Scratch","description":"Self-Attention from Scratch: Attention is a lookup table where every word asks \"who matters to me?\" - and learns the answer. Implement scaled dot-product…","url":"https://aiengineeringfromscratch.com/lesson?path=phases%2F07-transformers-deep-dive%2F02-self-attention-from-scratch","sourceUrl":"https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/02-self-attention-from-scratch/docs/en.md"},{"path":"phases/07-transformers-deep-dive/03-multi-head-attention","kind":"lesson","title":"Multi-Head Attention","description":"Multi-Head Attention: One attention head learns one relation at a time. Eight heads learn eight. Heads are free. Take more of them.","url":"https://aiengineeringfromscratch.com/lesson?path=phases%2F07-transformers-deep-dive%2F03-multi-head-attention","sourceUrl":"https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/03-multi-head-attention/docs/en.md"},{"path":"phases/07-transformers-deep-dive/04-positional-encoding","kind":"lesson","title":"Positional Encoding — Sinusoidal, RoPE, ALiBi","description":"Positional Encoding — Sinusoidal, RoPE, ALiBi: Attention is permutation-invariant. \"The cat sat on the mat\" and \"mat the on sat cat the\" produce the same…","url":"https://aiengineeringfromscratch.com/lesson?path=phases%2F07-transformers-deep-dive%2F04-positional-encoding","sourceUrl":"https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/04-positional-encoding/docs/en.md"},{"path":"phases/07-transformers-deep-dive/05-full-transformer","kind":"lesson","title":"The Full Transformer — Encoder + Decoder","description":"The Full Transformer — Encoder + Decoder: Attention is the star. Everything else — residuals, normalization, feed-forward, cross-attention — is the…","url":"https://aiengineeringfromscratch.com/lesson?path=phases%2F07-transformers-deep-dive%2F05-full-transformer","sourceUrl":"https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/05-full-transformer/docs/en.md"},{"path":"phases/07-transformers-deep-dive/12-kv-cache-flash-attention","kind":"lesson","title":"KV Cache, Flash Attention & Inference Optimization","description":"KV Cache, Flash Attention & Inference Optimization: Training is parallel and FLOP-bound. Inference is serial and memory-bound. Different bottleneck, different…","url":"https://aiengineeringfromscratch.com/lesson?path=phases%2F07-transformers-deep-dive%2F12-kv-cache-flash-attention","sourceUrl":"https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/12-kv-cache-flash-attention/docs/en.md"}],"total":24,"offset":0,"limit":10,"nextOffset":10}
