Large Language Models (LLMs)
What a large language model is, how it produces text token by token, how it is trained, and what it is and is not good at.
Prompt Engineering Basics
How to write prompts that give language models clear instructions, context, examples, output contracts, and measurable success criteria.
Reasoning Models & Test-Time Compute
How reasoning models spend extra inference-time tokens, how provider controls differ, and when the extra cost and latency are worth it.
AI Agents
What makes an LLM system an agent, how tool use works, the canonical multi-agent patterns, and the MCP and A2A protocols that connect agents to tools and to each other.
RAG (Retrieval-Augmented Generation)
How retrieval-augmented generation grounds an LLM in external knowledge using embeddings and vector databases, how it compares to fine-tuning, and the production levers that make it work.
Embeddings Deep Dive
How embedding models work, the 2026 model landscape, how to choose one, dimensions and Matryoshka, vector quantization, domain adaptation, hybrid retrieval, and the production pitfalls that quietly halve recall.
Context & Prompt Engineering
The context window as a finite budget, why prompt engineering grew into context engineering, context rot, and the long-horizon techniques - compaction, structured note-taking, sub-agents, and just-in-time retrieval.
Tooling and Frameworks
A map of the AI application tooling landscape -- orchestration frameworks, connectivity protocols, vector databases, evaluation and observability, and the LLMOps discipline that ties them together.
MDN MCP Server
Connect AI tools and IDEs to MDN's live documentation and browser compatibility data via the official MDN MCP server.
MCP & A2A in Production
Production notes for Model Context Protocol and Agent2Agent - architecture, transports, authorization, security, registries, observability, and adoption trade-offs.
Evaluation & LLMOps
How to test non-deterministic LLM systems with datasets, scorers, and LLM-as-judge; eval-driven development and harness engineering; and the LLMOps discipline of operating prompts, models, and agents in production.
Eval Datasets & Synthetic Data
How to build golden eval datasets for LLM systems, label them consistently, split them safely, and use synthetic data without fooling yourself.
AI Safety & Guardrails
What guardrails are and how they work, their documented limitations, the attack surface (prompt injection and jailbreaking), red-teaming as an evaluation method, and the layered, defense-in-depth approach to deploying LLMs responsibly.
Agent Security & Sandboxing
How to secure tool-using LLM agents with threat modeling, prompt-injection defenses, sandboxing, egress controls, least-privilege tools, approval gates, and audit trails.
Knowledge Management with LLMs
The design axis behind RAG, just-in-time retrieval, structured note-taking, the LLM-wiki pattern, and llms.txt - where synthesized knowledge lives, who maintains it, and when to use each.
AI-Assisted Software Development
Using LLM coding agents inside an engineering workflow - the alignment-before-generation methodologies (SPDD, architect-as-orchestrator), architecture patterns that suit AI (deep modules, vertical slices), evidence on productivity, and the economics driving adoption.
Agent Skills
Portable, on-demand workflow packages for coding agents - what they are, how they differ from rules and project memory, how to use them across tools, and how to author your own with examples.
Choosing a Model & Reading Benchmarks
How to choose an LLM for a product task, read benchmarks without being fooled, build a small internal eval set, and manage model upgrades safely.
Cost, Latency & Model Routing
Token economics, latency drivers, and practical patterns for choosing model tiers, caching, and fallback chains - without treating cost as an afterthought.
Structured Outputs
Getting reliable JSON and schema-bound responses from LLMs - native structured output modes, validation and repair loops, and when structure beats free-form prose.
Fine-Tuning & Model Adaptation
How to choose between prompting, RAG, supervised fine-tuning, preference tuning, LoRA, QLoRA, distillation, and continued pretraining without using the wrong tool for knowledge injection.
Multimodal & Voice
How multimodal models handle images, documents, audio, speech, and realtime voice agents, with practical architecture, latency, retrieval, evaluation, safety, and privacy guidance.
AI in Products
Product and UX patterns for shipping LLM features to users - interaction models, streaming, trust, when not to use AI, and graceful degradation.
AI Search, GEO & llms.txt
How AI answer engines retrieve and cite web content, what GEO and AEO honestly mean, how AI crawlers identify themselves, where llms.txt fits, and how to make documentation crawlable without chasing myths.
Privacy & Data Handling
Practical guidance for builders on PII in prompts, retention and residency, logging risks, and minimization patterns when shipping LLM features.
AI Regulation for Builders
A practical, builder-focused snapshot of AI regulation, including the EU AI Act, GDPR interplay, NIST AI RMF, ISO/IEC 42001, and implementation checklists.
Project Memory & Rules
Always-on context for coding agents - AGENTS.md, CLAUDE.md, GitHub Copilot instructions, Cursor rules, Gemini context, Windsurf rules, and how to split conventions from on-demand skills.
Human-in-the-Loop
Approval gates, escalation, and accountability patterns for agents and LLM features that act in the real world - maker-checker, confidence thresholds, and audit trails.
Which Pattern When?
A decision guide for the AI section - pick the right combination of LLM, RAG, agents, skills, structured outputs, and human review for common goals without rereading every page.
Debugging LLM Apps
Production troubleshooting for LLM features - classify the failure, inspect prompts retrieval tools and logs, and fix the right layer without guessing.
Cloud vs Local Models
When to use enterprise cloud AI platforms (Bedrock, SageMaker, Azure / Microsoft Foundry, Vertex AI) versus running open-weights models locally with Ollama, LM Studio, llama.cpp, or vLLM - with a decision framework and trade-off tables.
Serving LLMs at Scale
How production LLM serving works, covering prefill, decode, KV cache sizing, batching, prefix reuse, speculative decoding, quantization, parallelism, engines, autoscaling, and benchmarks.
Build a Local LLM App
A hands-on guide to running a model locally with Ollama or LM Studio and talking to it from your own simple app - with copy-paste code for a browser app, a Node.js CLI, and the OpenAI SDK.
Local & offline Copilot alternative
How to run a local coding assistant with Ollama, Continue, and Open WebUI, including model roles, reranking, context limits, and security notes.
AI Agent Harnesses
The AI coding-agent harnesses I run in the terminal -- opencode, pi, and herdr -- what a harness is, how they differ, and how I use them day to day.
AI Glossary
Short, plain-English definitions of the core AI, LLM, agent, and RAG terms used across this section, with links to the deep-dive pages.