Playbooks

CCAR-F: Claude Certified Architect Foundations — 2026 Blueprint

An exam-day reference for the CCAR-F certification covering all five domains, the six exam scenarios, anti-patterns, trade-offs, and scenario triggers.


Everything I’d want a teammate to walk into the Claude Certified Architect exam knowing — the five domains, the scenarios that trip people up, and the reasoning behind each answer.

QuestionsDurationFeePass markScenariosProctored
60120 min$125 USD720 / 10004 of 6Yes, closed-book

Exam Format & Domains

Launched March 12, 2026, Claude Certified Architect — Foundations (CCAR-F) was Anthropic’s first proctored technical certification. (Early coverage widely wrote it CCA-F; CCAR-F is the official code, and it now sits alongside an Associate, a Developer, and a Professional-tier Architect exam.) It validates that you can design and ship production-grade Claude applications at enterprise scale. Every question is scenario-based — a realistic production system with a problem, and you pick the architecturally correct fix among plausible alternatives. No simple recall questions.

Before you plan around this: the certification is offered through the Claude Partner Network — registration runs via the Anthropic Partner Academy, which validates partner credentials at login. Confirm your organization’s eligibility before budgeting exam time.

Five domains & weighting

DomainWeightCore concepts
1. Agentic Architecture & Orchestration27%Agentic loop, stop_reason, hub-and-spoke, subagents, hooks, task decomposition, session state
2. Tool Design & MCP Integration18%MCP servers (resources/tools/prompts), JSON Schema, stdio vs SSE, tool descriptions, tool_choice
3. Claude Code Configuration & Workflows20%CLAUDE.md hierarchy, skills, slash commands, plan mode, path-specific rules, CI/CD (-p flag)
4. Prompt Engineering & Structured Output20%PRECISE (community mnemonic), few-shot, XML tags, JSON schemas, tool_use for extraction, validation-retry, Batch API
5. Context Management & Reliability15%Context window, prompt caching, token budgets, CALM (informal study aid), escalation, error propagation, human-in-the-loop

Scoring: Scaled 100–1000, pass at 720. Domain-weighted — you can’t pass by acing one domain and ignoring others. 4 of 6 published scenarios are randomly selected per exam; each provides the context for ~15 questions.

The 6 Exam Scenarios

You get 4 of these 6 on exam day — randomly selected. Study all six so you’re not caught off guard. Each is a realistic production system that spans multiple domains.

Scenario 1 — Customer Support Resolution Agent

  • What it tests: Agent SDK + MCP tools + escalation triggers + error propagation
  • Key decisions: When to escalate to a human, how to structure the agentic loop, how to handle policy gaps and ambiguity

Scenario 2 — Code Generation with Claude Code

  • What it tests: CLAUDE.md config + plan mode + slash commands + built-in tools
  • Key decisions: Autonomous PR review/generation, skills with context: fork, how to structure project instructions

Scenario 3 — Multi-Agent Research System

  • What it tests: Coordinator-subagent orchestration + context isolation + structured passing
  • Key decisions: Hub-and-spoke vs pipeline, parallel vs sequential, preventing context pollution

Scenario 4 — Developer Productivity with Claude

  • What it tests: Built-in tools + MCP servers + tool descriptions + tool distribution
  • Key decisions: When to use existing MCP servers vs build custom, IDE integration patterns

Scenario 5 — Claude Code for CI/CD

  • What it tests: Non-interactive mode (-p flag) + structured output + independent review
  • Key decisions: Headless pipeline design, --output-format json, when to use --json-schema

Scenario 6 — Structured Data Extraction

  • What it tests: JSON schemas + tool_use for extraction + validation-retry loops + batch processing
  • Key decisions: Nullable fields to reduce hallucination, single-pass vs multi-pass, Message Batches API

§1 — Agentic Architecture & Orchestration (27%)

The heaviest domain and the one most candidates find hardest. It covers the full lifecycle of the agentic loop, multi-agent coordination, hooks, task decomposition, and session management. What follows is scoped to what the exam asks; for the same patterns treated as design decisions rather than answers, see the AI architecture field guide.

The agentic loop — the single most important concept

The agentic loop lifecycle:

1. Send request (history + tools)
   → 2. Receive response
   → 3. Check stop_reason        ← the key step
   → 4. Execute tools (if tool_use)
   → 5. Append results (to history)
   → Loop back to 1
  • stop_reason === "tool_use" → Claude wants to call tools → execute them → append results → loop back.
  • stop_reason === "end_turn" → Claude is finished → terminate the loop.
  • stop_reason === "pause_turn" → a long server-side-tool turn (web search, code execution) checkpointed instead of finishing → re-send the conversation unchanged to resume. It’s a checkpoint, not a termination — don’t treat it as “done.”
  • The stop_reason field is a structured API signal — the only reliable termination mechanism.

Anti-patterns (exam favorites):

  • Parsing natural language to detect “done” / “task complete” — non-deterministic, fails silently.
  • Arbitrary iteration cap as the primary stopping mechanism — fragile, not semantically meaningful.
  • Checking assistant text content like response.content[0].text.includes('complete') — brittle string matching.

Multi-agent orchestration — hub-and-spoke

The hub-and-spoke (coordinator-subagent) pattern:

            Coordinator (hub)
           /       |        \
  Web search   Doc analysis   Synthesis
   subagent      subagent      subagent
  • The coordinator manages all inter-subagent communication, error handling, and routing. Subagents never talk to each other directly.
  • Context isolation: each subagent gets only the context it needs — never the coordinator’s full context (prevents context pollution).
  • Structured passing: when handing findings between subagents, pass structured data (JSON with source, date, content), not plain text blobs — preserves attribution.
  • Minimal footprint: subagents should be scoped to a single responsibility and return a focused result.

Anti-pattern: Sharing the coordinator’s full conversation history with every subagent. This causes context pollution — subagents get confused by irrelevant context and waste tokens.

Agent SDK hooks

Hooks are programmatic enforcement points — Python/TypeScript functions invoked by the Agent SDK at specific loop points. They are not invoked by Claude (the model never sees them).

Hook typeWhen it firesUse for
Pre-toolBefore a tool is executedValidate inputs, check permissions, block dangerous actions
Post-toolAfter a tool returnsLog results, sanitize output, enforce PII rules
Pre-messageBefore Claude’s response is sent to the userContent filtering, compliance checks, redaction

Trade-off — hooks vs prompt-based guardrails: Hooks = programmatic, deterministic, enforced by code, can’t be jailbroken → use for hard constraints (compliance, PII blocking, permission checks). Prompt instructions = flexible, natural-language guidance, can be circumvented → use for soft guidelines (tone, style, preference). Default to hooks for anything safety-critical.

Task decomposition & execution patterns

PatternWhen to useTrade-off
SequentialSteps depend on prior results; order mattersSlower but safer; each step informed by the last
ParallelIndependent subtasks with no dependenciesFaster but costs more tokens; harder to debug
Dynamic planningTask scope unclear upfront; agent decides next stepsMost flexible but needs guardrails to prevent drift

Session management: Know: fork_session creates an isolated branch (subagent work doesn’t pollute the main conversation). Sessions can be resumed with conversation history. State can be in-context (conversation history) or external (database/file); external is more durable for long-running agents.

§2 — Tool Design & MCP Integration (18%)

Tests your ability to design tool interfaces, write schemas Claude can reliably use, and build/integrate MCP servers.

MCP (Model Context Protocol) — three primitives

PrimitiveWhat it isAnalog
ToolsFunctions Claude can call — input schema, returns a resultPOST / RPC
ResourcesData Claude can read — files, DB records, URIsGET / read-only
PromptsReusable prompt templates exposed by the serverTemplates

Tool schema design

  • JSON Schema defines tool inputs. Every tool needs name, description, and input_schema.
  • Descriptions are how Claude decides which tool to use. They must be precise, use distinct verbs, and have narrow scope. “Two good matches” and “no good match” are both failures of description quality.
  • Nullable fields (e.g. "type": ["string", "null"]) reduce hallucination — Claude can explicitly return null instead of inventing a value.

Anti-patterns:

  • Vague descriptions (“does stuff with data”) — Claude can’t select the right tool.
  • Overlapping tools with similar descriptions — Claude oscillates between them.
  • Missing required fields in the JSON Schema — Claude hallucinates the structure.
  • Giant monolithic tools that do many things — hard to select, hard to debug.

Transport: stdio vs SSE

stdio (standard I/O):

  • Subprocess on the same machine
  • Lower latency, simpler setup
  • Best for local development, CLI tools

SSE (Server-Sent Events):

  • HTTP-based, works across network boundaries
  • Supports real-time streaming
  • Best for remote/cloud servers, team sharing
  • Note: the standalone HTTP+SSE transport is deprecated in favor of Streamable HTTP, which is now the recommended remote transport. Treat “SSE” here as shorthand for HTTP-based remote transport.

Scenario trigger: “Stream large file contents across a network” → SSE. “Local subprocess, same machine” → stdio. “Share MCP server across a team” → SSE (or Streamable HTTP).

tool_choice parameter

ValueBehaviorUse when
autoClaude decides whether/which tool to callDefault; most scenarios
anyClaude must call a tool (any one)Force tool use; extraction pipelines
tool (specific)Claude must call this exact toolDeterministic extraction; structured output
noneClaude cannot use any toolsForce text-only response

Structured error responses

  • Transient errors (network timeout, rate limit) → retry with backoff.
  • Business errors (record not found, invalid input) → return structured error to Claude so it can reason about next steps.
  • Permission errors → escalate or inform the user; don’t retry silently.

§3 — Claude Code Configuration & Workflows (20%)

Tests your ability to configure Claude Code (Anthropic’s agentic CLI/IDE tool) for production projects and team workflows.

CLAUDE.md hierarchy & scoping

LevelLocationScope
EnterpriseSet by admin (operator)All users in the org — hard constraints
User / global~/.claude/CLAUDE.mdPersonal preferences across all projects
Project./CLAUDE.md (repo root)Shared team conventions, committed to VCS
Path-specific rules.claude/rules/*.md with YAML frontmatterActivated only when editing matching file paths
  • Higher levels override lower: Enterprise > User > Project.
  • An operator (enterprise admin) can set constraints that users cannot override — this is the trust hierarchy.
  • .claude/rules/ with path-scoped YAML frontmatter is the recommended approach over a monolithic CLAUDE.md — reduces irrelevant context.

Skills, commands & built-in tools

ConceptWhat it is
Custom slash commandsDefined in .claude/commands/; reusable workflows invoked by /command-name
Skills (SKILL.md)Project- or user-scoped instructions for specific tasks (e.g. “how to create a docx”); context: fork in frontmatter runs the skill in an isolated sub-agent
Built-in toolsRead, Write, Edit, Bash, Search, List — Claude Code’s native file/code tools
MCP servers in Claude CodeConfigured in .claude/mcp.json (project) or ~/.claude/mcp.json (user)

Scenario trigger: context: fork = the skill runs as an isolated sub-agent, preventing verbose output from polluting the main session. Use it for tasks that generate large output (code generation, data processing).

Execution modes

ModeHow it worksUse when
Plan modeClaude plans the approach before executing; shows the plan for reviewComplex tasks; want to review before committing
Direct executionClaude executes immediatelySimple / well-understood tasks
Non-interactive (-p flag)Headless mode, no user prompts; outputs to stdoutCI/CD pipelines; add --output-format json for structured output

CI/CD (scenario 5): For CI/CD: use claude -p "review this PR" --output-format json --json-schema schema.json. This gives deterministic, machine-parseable output. Remember: -p = non-interactive / print mode.

§4 — Prompt Engineering & Structured Output (20%)

The PRECISE mnemonic

A community mnemonic — not official Anthropic terminology — for a structured approach to designing system prompts. Know the acronym and what each component does:

LetterComponentWhat it does
PPersonaWho Claude is (“You are a senior code reviewer”)
RRoleThe function it performs
EExplicit instructionsClear, direct guidance
CContextBackground information
IInstructionsStep-by-step task rules
SStepsOrdered process to follow
EExamplesFew-shot demonstrations of ideal output

Key insight: The Examples component is the behavioral anchor — it shows Claude what a correct response looks like in terms of length, tone, and structure. Without examples, Claude interpolates from pre-training, producing variance. This is often the fastest fix for inconsistent output.

Core techniques

  • Few-shot prompting: 2–3 input/output examples to anchor behavior. Place them after the instructions, before the actual task.
  • XML tags for context structure: <document>, <instructions>, <example> — Claude respects these boundaries for parsing and retrieval.
  • Chain-of-thought / extended thinking: for complex reasoning, let Claude think step-by-step before producing the final answer.
  • Prefilled assistant responses: start the assistant turn with a partial response to steer format/structure. Note: most current frontier Claude models (Opus 4.6+, Sonnet 4.6+, Fable 5) reject a prefilled assistant turn — prefer schema-constrained structured outputs there.

Structured output via tool_use

The exam’s preferred pattern for extracting structured data: define a tool whose input_schema matches your desired JSON structure, then set tool_choice to force Claude to call it. This gives you validated, schema-conforming output.

Trade-off — tool_use extraction vs raw JSON prompting: tool_use with JSON Schema = schema-validated, deterministic structure, nullable fields reduce hallucination → preferred for production. Prompting for JSON = simpler to set up, but output can drift from schema, requires post-processing validation → acceptable for prototyping only.

Validation-retry loops

  • After extraction, validate the output programmatically against the schema.
  • If validation fails, send the error back to Claude with the original input and ask it to fix the specific issue.
  • Cap retries (2–3) to avoid infinite loops on truly malformed input.

Batch processing — Message Batches API

  • For high-volume extraction (hundreds/thousands of documents), use the Message Batches API — asynchronous, 50% cheaper than real-time, 24-hour processing window.
  • Trade-off: latency (not real-time) for cost savings — ideal for overnight/batch ETL.
  • Per-item error isolation: if one item in a batch fails, the rest still succeed. Design for this.

§5 — Context Management & Reliability (15%)

The lightest domain by weight but consistently underestimated. These questions are often the easiest marks available once you know the patterns — and the easiest to lose if you don’t.

Context window fundamentals

  • Context window = input tokens + output tokens. Know the model limits: Opus (4.6/4.7/4.8), Sonnet (4.6/5), and Fable 5 = 1M-token context; only Haiku 4.5 = 200K.
  • System prompt + conversation history + tool definitions + tool results all consume input tokens. They add up fast in agentic loops.
  • Progressive summarization: as conversation grows, periodically summarize older turns and replace the raw history → keeps context under budget without losing critical information.
  • Conversation compaction: similar idea, applied at the system level — compress older context to make room for new.

Prompt caching

  • Use cache_control breakpoints on large, static content blocks (system prompts, reference docs, tool definitions) that don’t change between turns.
  • Cached reads are 90% cheaper than uncached — massive savings in multi-turn agentic loops where the system prompt is repeated every turn.
  • Cache has a 5-minute TTL by default, with a 1-hour TTL option also available — if the next request comes within the TTL window, you get the cached price.
  • Trade-off: write cost to populate cache is 25% more than a normal read, so caching only pays off if you’re making multiple requests against the same prefix.

Trade-off — when caching pays off: Cache if: multi-turn conversation, repeated system prompt, agentic loop (many iterations). Don’t cache if: one-shot request with unique content. The break-even is roughly 2+ requests with the same cached prefix within 5 minutes.

The CALM mnemonic

Context-Aware LLM Management — an informal study aid, not an official Anthropic concept — for managing what goes into the context window:

  • Curate: only include information Claude actually needs for this turn.
  • Arrange: put the most important information at the start and end of context (primacy/recency bias).
  • Limit: enforce token budgets; summarize or truncate when approaching limits.
  • Monitor: track token usage, cache hit rates, and quality over time.

Error handling & escalation

Error typePattern
Tool failure (transient)Retry with exponential backoff; return structured error to Claude if retries exhausted
Reasoning failureSelf-evaluation / confidence scoring; if confidence below threshold → escalate
Environment failureGraceful degradation; fall back to cached/stale data or inform user
Policy gap / ambiguityHuman-in-the-loop escalation — don’t guess, ask

Anti-pattern: Using self-reported confidence scores as the sole escalation signal. Claude’s self-assessed confidence is unreliable — it may be confidently wrong. Combine with programmatic validation (schema checks, business-rule assertions) for robust escalation.

Model Selection & API Essentials

ModelStrengthUse when
Opus (largest)Highest reasoning, complex tasks, nuanced judgmentHard reasoning, complex orchestration, low-volume high-stakes
Sonnet (mid)Best balance of capability and cost; fastDefault for most production workloads, agentic loops, coding
Haiku (smallest)Fastest, cheapest, good for simple tasksClassification, routing, simple extraction, high-volume low-complexity

Trade-off — model selection: Bigger = better reasoning but slower and costlier. Right-size to the task: use Haiku for routing/classification, Sonnet for the main workload, Opus for the hardest reasoning steps. In multi-agent systems, subagents can use smaller models than the coordinator.

API concepts the exam assumes

  • Messages API: the core endpoint. Send a list of messages (system, user, assistant), receive a response with content blocks and stop_reason.
  • Streaming: stream: true for real-time token delivery. Use for user-facing responses.
  • Extended thinking: thinking blocks that let Claude reason before responding. On current models this is adaptive (thinking: {type: "adaptive"} plus an effort level) rather than a fixed token budget. Useful for complex tasks; costs extra tokens.
  • System / Operator / User hierarchy: operator instructions (set by the app developer) can restrict what user prompts can override. Users can’t elevate their own permissions beyond what the operator allows.

Anti-Pattern Master List

The exam loves presenting anti-patterns as plausible answer choices. Memorize these — they’re the wrong answers that look right.

Anti-patternWhy it’s wrongCorrect approach
Parse NL to detect loop endNon-deterministic; Claude may paraphraseCheck stop_reason === "end_turn"
Arbitrary iteration cap as primary stopNot semantically meaningfulUse stop_reason; iteration cap as safety net only
Share full coordinator context with subagentsContext pollution; wasted tokensPass only relevant, structured context
Vague / overlapping tool descriptionsClaude misroutes tool callsDistinct verbs, narrow scope, precise descriptions
Self-reported confidence as sole escalation triggerClaude can be confidently wrongCombine with programmatic validation
Prompt-only guardrails for hard safety constraintsCan be jailbroken / circumventedUse hooks for hard constraints
Monolithic CLAUDE.md with all rulesIrrelevant context for most edits.claude/rules/ with path-specific scoping
JSON via prompting alone in productionOutput drifts from schematool_use with JSON Schema for structured extraction
Retry on all errors equallyWastes tokens on permanent failuresClassify errors: transient → retry; business → return; permission → escalate
No token budget monitoringSilent degradation as context growsTrack usage; progressive summarization; prompt caching

Trade-off Master List

Decisions the exam tests repeatedly. Know the trigger phrase → right answer.

DecisionOption AOption BPick A whenPick B when
Loop terminationstop_reasonIteration capAlways primarySafety net only
GuardrailsHooks (code)Prompt instructionsHard safety constraintsSoft style/tone guidance
ExecutionSequentialParallelSteps depend on prior resultsIndependent subtasks
TransportstdioSSELocal, same machineRemote, team sharing, streaming
Structured outputtool_use + schemaPrompt for JSONProduction extractionQuick prototyping only
Model sizeHaikuOpusSimple / routing / high-volumeComplex reasoning / high-stakes
CachingCache (90% cheaper reads)No cacheMulti-turn / agentic loopOne-shot unique requests
Batch vs real-timeBatches API (50% off)Real-time MessagesBulk ETL, overnight jobsUser-facing, low-latency
Context strategyProgressive summarizationFull historyLong conversations / budgetsShort conversations / precision
CLAUDE.md structurePath-scoped rulesMonolithic fileLarge repos, mixed stacksTiny projects
Execution modePlan modeDirect executionComplex / high-risk changesSimple / well-understood tasks
Skill isolationcontext: forkInline executionVerbose output / risk of pollutionSimple, brief tasks

Exam-Day Tips

  • Study all 6 scenarios — you get 4 randomly. Map each to its primary domains: Customer Support → agentic loop + escalation; Code Gen → Claude Code config; Multi-Agent → hub-and-spoke + context isolation; CI/CD → -p flag + structured output; Data Extraction → tool_use + validation-retry.
  • stop_reason is the answer for any loop-termination question. If an option says “parse the text for ‘done’” — eliminate it.
  • Two answer choices will “work” — pick the architecturally superior one. The exam rewards the production-grade, scalable, safe choice over the quick hack.
  • Hooks for hard constraints, prompts for soft guidance. This distinction appears across multiple scenarios.
  • Context isolation is a theme. Any time a subagent gets “too much” context, or a coordinator shares everything, it’s the wrong answer.
  • Structured > unstructured. Passing structured JSON between agents beats plain text. tool_use extraction beats prompting for JSON. Schema-validated output beats unvalidated.
  • Right-size the model. Using Opus for classification or Haiku for complex reasoning are both wrong in scenario questions.
  • Prompt caching math: 90% cheaper reads, 25% more expensive writes, 5-min default TTL (1-hour option available), break-even at ~2 requests. Know this for cost-optimization questions.
  • 2 min/question average. Don’t overthink — recognize the pattern (it maps to one of the anti-patterns or trade-offs above), pick the answer, and move on. Flag and return if unsure.

Preparation resources (free): Anthropic Academy (anthropic.skilljar.com): 13+ free courses covering all domains. Key courses: Building Applications with the Claude API (8+ hrs), Claude Code in Action, Introduction to MCP, AI Fluency Framework. Also: the official Exam Guide PDF (12 sample questions with explanations) and the official 60-question practice exam on Skilljar. Score 850+ on the practice before booking the real exam. Note the three-way split as of mid-2026: the public Academy hosts the free training courses, registration goes through the Partner Academy (anthropic-partners.skilljar.com, partner login required), and the exam itself is scheduled and proctored via Pearson VUE (OnVUE). Retakes are capped at 4 attempts per rolling 12 months, with waiting periods that lengthen after each failure (14 days, then 30, then 90).


Claude Certified Architect — Foundations (CCAR-F), 2026 Blueprint. Built from the official exam guide, Anthropic Academy materials, and community sources. Independent study aid — not affiliated with or endorsed by Anthropic. Exam format, fees, and scheduling change; verify current details on Pearson VUE’s Anthropic certification page before booking.

Riddam Jain

Staff Engineer · Amsterdam

I write about cloud and application architecture, AI, and leading engineers. If this was useful, let's connect.