These documents are self-study notes compiled for personal learning in applied artificial intelligence. The material is intentionally crude, austere, and simple. I wrote them for my own reference, but glad if somebody else finds them useful.
_Article
Applied AI Engineering Notes
Track 1 — Product Strategy for AI
Module 1 — AI Decision Framework
1.1. When NOT to Use AI
- 1.1.1 Deterministic logic suffices
- 1.1.2 Tail risk too high (legal, safety)
- 1.1.3 Insufficient or poor-quality data
1.2. Identifying High-Leverage AI Opportunities
- 1.2.1 Repetitive cognitive work
- 1.2.2 Bottleneck workflows
- 1.2.3 Data-rich, judgment-light tasks
1.3. Deterministic Systems vs AI Systems
- 1.3.1 When determinism wins (correctness, audit)
- 1.3.2 When AI wins (open-ended, ambiguous input)
- 1.3.3 Hybrid: deterministic shell + AI step
1.4. Build vs Buy vs Embed AI
- 1.4.1 Buy when commodity
- 1.4.2 Build when defensibility lives in workflow
- 1.4.3 Vendor lock-in & switching cost
1.5. AI Feasibility vs Business Value
- 1.5.1 Feasibility axis: data, latency, accuracy, cost
- 1.5.2 Value axis: revenue, cost saved, risk reduced
- 1.5.3 Prioritization matrix
Module 2 — AI-Native Product Design
2.1. UX for AI Products
- 2.1.1 Input affordances (text, voice, file)
- 2.1.2 Output presentation (streaming, structured)
- 2.1.3 Error and retry states
2.2. AI-Native Product Design
- 2.2.1 Build from AI primitives, not retrofit
- 2.2.2 Conversational vs task-oriented surfaces
- 2.2.3 Memory & continuity across sessions
2.3. Designing Around Uncertainty
- 2.3.1 Surfacing confidence
- 2.3.2 Retraction & correction UX
- 2.3.3 Source attribution
2.4. Latency, Feedback, and User Perception
- 2.4.1 Streaming to mask wait time
- 2.4.2 "Thinking" indicators with substance
- 2.4.3 Latency budgets per surface
2.5. Invisible AI vs Explicit AI UX
- 2.5.1 Ambient AI (autocomplete, ranking)
- 2.5.2 Explicit AI (chat, agent)
- 2.5.3 When to disclose AI involvement
Module 3 — Trust & Human Interaction
3.1. Human Trust Models
- 3.1.1 Calibration: confidence vs accuracy
- 3.1.2 Over-trust (automation bias)
- 3.1.3 Under-trust (algorithm aversion)
3.2. Confirmation UX
- 3.2.1 Reversibility ladder
- 3.2.2 Blast radius before approval
- 3.2.3 Batch vs per-action confirmation
3.3. Human-in-the-Loop Design
- 3.3.1 Review-queue patterns
- 3.3.2 Sampling strategies for review
- 3.3.3 Active-learning loops
3.4. Escalation Paths
- 3.4.1 Confidence-based handoff
- 3.4.2 Human fallback for failure
- 3.4.3 Context-logging for the human
3.5. Permission Boundaries and User Control
- 3.5.1 Per-action vs blanket consent
- 3.5.2 Scoped credentials per agent
- 3.5.3 Kill switch & global pause
Module 4 — Workflow Transformation
4.1. Workflow Redesign
- 4.1.1 Map current process before redesign
- 4.1.2 Identify decision points vs execution
- 4.1.3 Eliminate steps before automating
4.2. Replacing vs Assisting Human Work
- 4.2.1 Augmentation: speed and quality lift
- 4.2.2 Automation: full task removal
- 4.2.3 Hybrid: AI drafts, human approves
4.3. Operational Bottlenecks and AI Leverage
- 4.3.1 Find the queue with the longest wait
- 4.3.2 Convert serial work to parallel
- 4.3.3 Throughput per FTE before/after
4.4. Process Decomposition for AI Integration
- 4.4.1 Atomic-task vs end-to-end automation
- 4.4.2 Tool-callable subtasks
- 4.4.3 Idempotent units of work
4.5. Organizational Readiness
- 4.5.1 Enterprise AI readiness
- 4.5.2 Data hygiene baseline
- 4.5.3 Change capacity & fatigue
- 4.5.4 Sponsor & champion presence
Module 5 — Business Models & Economics
5.1. ROI Models
- 5.1.1 Cost per task before/after
- 5.1.2 Quality lift (revenue, retention)
- 5.1.3 Substitution rate (% of human work replaced)
5.2. Pricing AI Products
- 5.2.1 Seat-based
- 5.2.2 Usage-based (tokens, calls)
- 5.2.3 Outcome-based
5.3. Cost Structures of AI Products
- 5.3.1 Variable cost per call (model COGS)
- 5.3.2 Fixed inference infra
- 5.3.3 Training & eval costs
5.4. Usage-Based Pricing vs Seat-Based Pricing
- 5.4.1 Predictability for the customer
- 5.4.2 Margin protection for the vendor
- 5.4.3 Incentive alignment with customer value
5.5. Margin Compression from Model Costs
- 5.5.1 COGS scales with usage
- 5.5.2 Caching, distillation, smaller models
- 5.5.3 Self-hosted alternatives
Module 6 — Enterprise Adoption Strategy
6.1. Adoption Barriers
- 6.1.1 Trust & accuracy concerns
- 6.1.2 IT / security gatekeeping
- 6.1.3 Legal & compliance review
6.2. Internal Adoption Strategy
- 6.2.1 Pilot-team selection
- 6.2.2 Champion network
- 6.2.3 Expansion playbook
6.3. Change Management for AI Systems
- 6.3.1 Training & enablement
- 6.3.2 Incentive alignment
- 6.3.3 Sunsetting old workflows
6.4. Stakeholder Alignment
- 6.4.1 Exec sponsorship
- 6.4.2 IT / security partnership
- 6.4.3 Operations & frontline buy-in
6.5. Executive Buy-In vs Operator Buy-In
- 6.5.1 Top-down narrative vs bottom-up traction
- 6.5.2 Operator veto power
- 6.5.3 Sequencing the two
Track 2 — Enterprise AI Architectures
Module 1 — System Types (Use Cases)
1.1. Retrieval & RAG Systems
- 1.1.1 Data Ingestion & Retrieval Pipeline
- 1.1.2 Document Ingestion Pipeline (PDFs, Tables, OCR)
- 1.1.3 Hybrid Search (Semantic + Keyword)
- 1.1.4 Vector Databases & Indexing
- 1.1.5 Reranking & Cross-Encoders
- 1.1.6 Context Construction & Grounding
- 1.1.7 Retrieval Evaluation (RAGAS, TruLens)
1.2. Traditional ML Systems
- 1.2.1 Supervised learning
- 1.2.2 Feature engineering
- 1.2.3 Training pipelines
- 1.2.4 Model serving
- 1.2.5 Retraining loops
- 1.2.6 Model monitoring
- 1.2.7 Fraud
- 1.2.8 Churn
- 1.2.9 Lead scoring
- 1.2.10 Forecasting
- 1.2.11 Anomaly detection
1.3. Fine-Tuning Systems
- 1.3.1 Supervised fine-tuning (SFT)
- 1.3.2 Preference optimization
- 1.3.3 Evaluation
- 1.3.4 Dataset quality
- 1.3.5 Labeling systems
1.4. Workflow / Agent Systems
- 1.4.1 State machines
- 1.4.2 Workflow engines
- 1.4.3 Retries
- 1.4.4 Idempotency
- 1.4.5 Approval gates
- 1.4.6 Tool execution
- 1.4.7 Auditability
- 1.4.8 Failure recovery
- 1.4.9 XState
- 1.4.10 Temporal
1.5. Recommendation Systems
- 1.5.1 Ranking models
- 1.5.2 Feedback loops
- 1.5.3 Candidate generation
- 1.5.4 Relevance systems
1.6. Knowledge Graph Systems
- 1.6.1 Entity modeling
- 1.6.2 Graph queries
- 1.6.3 Graph enrichment
- 1.6.4 Compliance applications
1.7. Central AI Platform
- 1.7.1 Ingestion
- 1.7.2 Permissions
- 1.7.3 Model routing
- 1.7.4 Observability
- 1.7.5 Evaluation
- 1.7.6 Governance
- 1.7.7 Tool access
- 1.7.8 Model abstraction
Module 2 — Deployment & Compliance Models
2.1. Isolation & Deployment Tiers
- 2.1.1 Cloud + External APIs (API-first)
- 2.1.2 VPC + Private Endpoints (Sovereign Cloud)
- 2.1.3 On-Prem + Air-Gapped (GPU Infra)
- 2.1.4 Hybrid Architectures
- 2.1.5 Vendor Strategy & Model Portability
2.2. Governance as Architecture
- 2.2.1 Audit-heavy architectures
- 2.2.2 IAM & Access Control
- 2.2.3 Cost Management & Token Budgets
- 2.2.4 Stakeholder Alignment & Architecture Politics
- 2.2.5 Trust-building & Phased Adoption
Module 3 — Strategic Risk & Governance
3.1. Discovery & Consulting Framework
- 3.1.1 What problem are we solving?
- 3.1.2 Where does work stop?
- 3.1.3 What data can leave the boundary?
- 3.1.4 Who approves exceptions?
- 3.1.5 What failure is expensive?
3.2. Systemic Risk Assessment
- 3.2.1 Hallucination & Success detection (Fundamentals)
- 3.2.2 Prompt & Tool security (Fundamentals)
- 3.2.3 Legal & Compliance Review
- 3.2.4 Business Impact of Failure
Module 4 — Example: Claude Enterprise AI
4.1. Claude Enterprise Architecture
- 4.1.1 Claude Enterprise AI
- 4.1.2 Claude Agent SDK
- 4.1.3 Hub-and-Spoke Patterns
- 4.1.4 Model Context Protocol (MCP)
- 4.1.5 Deterministic Enforcement (Hooks)
- 4.1.6 Validation-Retry Loops
4.2. The Claude Enterprise AI playbook
- 4.2.1 Claude Enterprise Playbook (FDE model)
- 4.2.2 Engagement Phases (Discovery to Scale)
- 4.2.3 FDE Technical Artifacts
- 4.2.4 White-Glove Support Models
- 4.2.5 Continuous Feedback Loops
- 4.2.6 E2E Example: Automated Loan File Auditor
- 4.2.7 E2E Example: Multi-Source Underwriting Assistant
Track 3 — Evaluation, Reliability & Operations
- 1 Eval frameworks
- 2 Benchmarking
- 3 Hallucination detection
- 4 Regression testing
- 5 Canary deployments
- 6 Guardrails
- 7 Observability
- 8 Failure analysis
- 9 Prompt testing
- 10 Production monitoring
- 11 Cost/quality tradeoffs
Track 4 — AI Security & Governance
- 1 Prompt injection
- 2 Tool abuse
- 3 Agent isolation
- 4 Least agency
- 5 Identity boundaries
- 6 Secret handling
- 7 Sandboxing
- 8 Audit trails
- 9 Data leakage
- 10 Tenant isolation
- 11 Compliance
- 12 Policy enforcement
Track 5 — Agentic Systems
Module 1 — Foundations (mandatory)
1.1. Distributed Systems Thinking
- 1.1.1 State machines
- 1.1.2 Queues
- 1.1.3 Retries
- 1.1.4 Idempotency
- 1.1.5 Timeouts
- 1.1.6 Failure recovery
- 1.1.7 Event-driven architecture
- 1.1.8 Workflow engines
- 1.1.9 Human approval gates
- 1.1.10 Observability
- 1.1.11 Designing Data-Intensive Applications
- 1.1.12 Temporal concepts
- 1.1.13 Apache Airflow basics
- 1.1.14 AWS Step Functions ideas
1.2. Planning + Decision Systems
- 1.2.1 Finite State Machines (FSM)
- 1.2.2 Hierarchical State Machines
- 1.2.3 Behavior Trees
- 1.2.4 Goal-oriented planning
- 1.2.5 PDDL (Planning Domain Definition Language)
- 1.2.6 HTN (Hierarchical Task Networks)
- 1.2.7 Rule engines
- 1.2.8 Constraint solving
- 1.2.9 XState
- 1.2.10 LangGraph
- 1.2.11 PDDL literature
- 1.2.12 Robotics planning basics
1.3. Retrieval + Grounding
- 1.3.1 Embeddings
- 1.3.2 Vector search
- 1.3.3 Hybrid search
- 1.3.4 Reranking
- 1.3.5 Chunking
- 1.3.6 Memory systems
- 1.3.7 Context construction
- 1.3.8 Grounded generation
- 1.3.9 Google Cloud Vertex AI retrieval patterns
- 1.3.10 Typesense
- 1.3.11 Pinecone
- 1.3.12 RAG (Retrieval-Augmented Generation)
Module 2 — LLM Systems
2.1. Tool Use + Execution
- 2.1.1 Tool calling
- 2.1.2 Structured outputs
- 2.1.3 Function schemas
- 2.1.4 MCP (Model Context Protocol)
- 2.1.5 A2A (Agent-to-Agent)
- 2.1.6 ADK patterns
- 2.1.7 Sandboxing
- 2.1.8 Permission boundaries
- 2.1.9 Anthropic MCP
- 2.1.10 Google ADK
- 2.1.11 CLI-agent patterns (Codex / Gemini / Claude Code)
2.2. Prompting for Systems
- 2.2.1 Prompt contracts
- 2.2.2 Tool instructions
- 2.2.3 Output guarantees
- 2.2.4 Failure handling
- 2.2.5 Repair loops
- 2.2.6 Self-checking
- 2.2.7 Reflection
2.3. Evaluation
- 2.3.1 Offline evals
- 2.3.2 Online evals
- 2.3.3 Task completion metrics
- 2.3.4 False success detection
- 2.3.5 Human review systems
- 2.3.6 Cost vs latency tradeoffs
Module 3 — Security (critical)
3.1. Agent Security
- 3.1.1 Prompt injection
- 3.1.2 Indirect prompt injection
- 3.1.3 Memory poisoning
- 3.1.4 Tool abuse
- 3.1.5 Credential leakage
- 3.1.6 Sandbox escape
- 3.1.7 Identity isolation
- 3.1.8 Approval boundaries
- 3.1.9 Least privilege
- 3.1.10 Audit logs
- 3.1.11 Anthropic Claude Code incidents
- 3.1.12 Prompt injection research
- 3.1.13 Agent exploit case studies
Module 4 — Production Architecture
4.1. Real Agent Architectures
- 4.1.1 Deterministic workflows + LLM steps
- 4.1.2 Single orchestrator + tools
- 4.1.3 Multi-agent systems
- 4.1.4 Event-driven workers
- 4.1.5 Human-in-the-loop systems
4.2. Programmable Agent Systems
- 4.2.1 Visual agent system designers architecture
- 4.2.2 Headless agent runtime
- 4.2.3 Pi Coding Agent
- 4.2.4 Commercial products
- 4.2.5 Open-source vs enterprise builders
- 4.2.6 Vendor lock-in and portability
4.3. Agent Harnesses
- 4.3.1 Custom harnesses
4.4. Agent Frameworks
- 4.4.1 Pi Mono
- 4.4.2 LangChain
- 4.4.3 LangGraph
- 4.4.4 Claude Agent SDK
- 4.4.5 OpenAI Agents SDK
- 4.4.6 Google ADK
- 4.4.7 Agent Frameworks Landscape
- 4.4.8 Declarative Compatibility
- 4.4.9 Enterprise Workflow Integration
- 4.4.10 The Brain & Guardrails Pattern
Track 6 — AI Infrastructure & Platform Engineering
- 1 Inference serving
- 2 Model gateways
- 3 Model routing
- 4 Cost optimization
- 5 Caching
- 6 Prompt management
- 7 Multi-model strategy
- 8 Vendor abstraction
- 9 Rate limits
- 10 Async pipelines
- 11 GPU economics
- 12 Batch vs realtime
Track 7 — Applied ML Systems (optional, advanced)
- 1 Fine-tuning
- 2 SFT (Supervised Fine-Tuning)
- 3 RLHF / RLAIF
- 4 Distillation
- 5 Training pipelines
- 6 Feature stores
- 7 Classical ML systems
- 8 Recommendation systems
- 9 Forecasting
- 10 Ranking systems