Anthropic interview, decoded.
Updated · Sources at the end
Anthropic's loop favours a practical work-sample style over LeetCode puzzles, paired with a dedicated values round that probes how you think about AI safety and responsible deployment. The company builds Claude, a family of large language models designed to be helpful, honest, and harmless, and it hires research engineers, machine learning and research scientists, software engineers, and product roles. It raised $65 billion at a $965 billion post-money valuation in its Series H in May 2026, and on 1 June 2026 it confidentially submitted a draft S-1 to the SEC, which gives it the option of an IPO; no listing date has been set.
Free · 2 minutes · no account
The questions for your exact Anthropic role and level.
Interviewing at Anthropic? Below are the questions candidates report. For the ones your exact role and level will get, paste the posting. Your first mock is free.
Who Anthropic hires
Anthropic hires research engineers, ML and research scientists, software and infrastructure engineers, and product roles, and it weighs demonstrated ability like open-source work, research, and technical writing over credentials, noting about half its technical staff had no prior ML experience and about half hold PhDs.
What is the Anthropic interview process?
Three to four weeks from recruiter call to decision, and faster with a competing offer. The shape is a recruiter screen that can fail you, the CodeSignal assessment, and an onsite of four to five hour-long sessions that ends with a values conversation. Team matching happens after the onsite. AI is not allowed in any live round or in the screen; the one take-home that permits it, the performance engineering test, says so explicitly.
Not a formality. Strong engineers fail it because they can say why AI and not why Anthropic. Expect questions on your understanding of what the company does, its B Corp status, and what you want next. Do not name a salary figure here.
One system built across escalating levels, Python standard library only, all the tests visible, the next level unlocking when the current one passes. Reported problems: an in-memory key-value store, a bank with transaction types, a file system, a package manager. Scored out of 600 with about 520 reported as the line, then an integrity review of the code. The section below goes through it.
One project in depth: the implementation, the decisions, the trade-offs, what you would change. Some managers add a code review across languages, reading a snippet and saying what it does and what is wrong with it.
Build from scratch and defend it: solve it cleanly with the standard library, then implement the data structure yourself; handle the edge cases; justify the complexity. Concurrency comes up across rounds rather than in one. Less code than at Meta or OpenAI, more reasoning per line.
Problems near Anthropic's own infrastructure: serving a large model efficiently, request batching, GPU utilisation, multi-region inference. Concurrency and multithreading again. The section below goes through the reported prompts.
How you think about AI safety, risk and deployment, and whether you can sit with an unresolved trade-off without reaching for a rehearsed line. Anthropic's recruiters say this is where most candidates fail.
The CodeSignal assessment: one system, four to six levels
Reported problems and how to solve themAnthropic's screen is CodeSignal's project-based format rather than its four-puzzle one: a single system that you build across levels, each level adding requirements on top of the last while every earlier method has to keep working. Ninety minutes, Python with only the standard library, the unit tests for the current level visible from the start, and the next level unlocking only when all of them pass. Reports differ on how many levels there are: four is what most candidates describe, including one who sat it in September 2026, while three independent accounts collected by one coach in June describe six, with the later waves invisible until you reach them. Prepare for six and be pleased with four. The spec is written loosely on purpose, so the tests are the real spec: run early, read what failed, adjust. Below are the reported problems and the shape each takes as it grows.
- In-memory key-value store: The most reported. Level one is set, get and delete; then filtered scans and range queries; then keys with a time to live; then something like persistence or a compaction step. Store a small object per key from the first line rather than a bare value, so timestamps and TTLs have somewhere to go without a rewrite.
- Bank with transaction types: Accounts and balances, then transfers, then a filtered transaction history, then interest or scheduled operations that depend on time. Still the most reported problem in 2026. It circulates on Blind and YouTube and interviewers know it, so a solution that looks memorised is a flag rather than a pass. Candidates disagree about how tidy to be: one who passed says he made no effort at all and shipped something close to a single function, while the coaching advice is to design for the levels you cannot see. The tests are what score, the review is what reads the code, and the resolution is the boring one: clean enough to extend at level four, not clean enough to win a code review.
- File system: Create and read, then permissions, then symlinks and mounts. The traps are circular references and inherited permissions; a tree of nodes with a resolve step you can extend beats path-string manipulation by level three.
- Merged entities and historical state: The hardest thing reported, and the one prepared candidates still miss. A late level asks you to merge two entities and keep both histories correct: which records the survivor now owns, what a query as of a date returns, whether an audit trail shows the merge or hides it, and what happens to identifiers other records point at. The candidate who scored 580 had built three practice systems and said this was the part his preparation had not covered. Design for it from level one by keeping events rather than only current values, and by letting an entity carry a pointer to the one it was merged into.
- Others reported: A package manager (install, dependencies, version constraints, conflicts), a build system (tasks, a DAG, caching, parallel runs), a text editor (insert and delete, undo and redo), a web crawler (fetch, parse, rate limit). Same rule: an interface at level one that has room for what level four will ask. Where a spec is ambiguous, the tests decide; one engineer described the whole screen as reading a loose spec and testing theories against a black-box grader.
- Scoring, and why a high one is not an offer of an interview: Six hundred points with later levels weighted more, and around 520 is the number coaching guides quote as the line. Treat it as a floor rather than a verdict. In a September 2026 thread one candidate completed all four levels for 580 and was rejected with no interview; a commenter reported 600 and the same outcome, another around 595 two years earlier, while a fourth advanced on 580. The screen filters, the profile decides, and the two happen in an order candidates cannot see. Two practical notes from the same thread: the in-test panel showed 95/100 while the portal showed 580/600 afterwards, which is the same result scaled and not a second mark; and a perfect score still goes through a review for code written to pass the tests rather than solve the problem, so meaningful names and a structure that reads as a design are the defence.
One candidate's timings, worth rehearsing against: level one done by minute 11, level two by 17, level three by 47, and the rest of the ninety on level four. Level three is where the hour goes, and level four is what you are buying time for. Spend the first ten to fifteen minutes on the data model and the method signatures even so. A class per entity, one method per action, storage separate from query logic: that is the whole strategy, because a level-one dictionary wrapper collapses when level three asks for a cross-cutting change. The preparation that candidate credits, and it is better than grinding problems: build three four-level systems yourself, a wallet or bank, a file system, an in-memory database. Write a specification for each level, implement it by hand, then have a model generate tests against your code with an explicit instruction not to modify or solve it. Fix what fails, then move up a level. That trains the actual skill, which is extending a system without breaking what already worked. Do it without AI writing the implementation, since none is allowed in the room and the review is built to catch it.
Anthropic system design interview: questions and how to answer them
Reported prompts and how to take themOne hour during the onsite, in a shared Google Doc instead of a drawing tool. The prompts sit close to Anthropic's own work: serving a large model, search over very large corpora, and the pipelines that feed a model. Interviewers look for distributed-systems fundamentals, where the bottleneck will be, which technology fits and why, and the trade-offs you accept, and they expect the parts specific to language models to be handled in detail. Concurrency comes up here too, as it does in the coding rounds.
- Design a Claude chat service: Start from the request path: authentication, rate limits per user and per organisation, a queue in front of the model servers, and tokens streamed back. Then spend most of the hour on the model tier: batching requests of very different lengths, GPU utilisation under bursty load, and where the conversation history lives and how much of it is resent on each turn.
- Serve a large language model efficiently through an API: The focus interviewing.io reports most. Cover request batching, queueing and GPU utilisation under variable load: continuous batching against fixed batches, keeping short requests fast when long ones share the batch, and what you autoscale on.
- Let a model handle several questions in a single thread: This is context management. Say how the thread's history is stored and trimmed to fit the context window, how questions that arrive close together are ordered, and what you cache, such as the shared start of the conversation, so each new turn does not pay for the whole history again.
- A distributed search system over a billion documents at a million queries a second, with LLM inference: Shard and replicate the index first, then place the model: run it on the small candidate set that retrieval returns, and give each stage its own latency budget so the model cannot eat the whole request.
- Hybrid search that combines keyword retrieval with semantic similarity: Two indexes, an inverted index and a vector index, queried in parallel and merged. Say how you combine the two scores, how you keep embeddings fresh when documents change, and what happens when the two disagree.
- An indexing pipeline for a billion documents, with concurrent crawling: Politeness limits and deduplication on the crawl side, a queue between crawling and indexing so either can fall behind, idempotent writes, and how you re-index everything when the embedding model changes.
- p95 latency jumped from 100 ms to 2,000 ms: find it and fix it: A debugging exercise inside the design round. Ask what changed first, a deploy, traffic or a dependency, then narrow it down stage by stage with metrics and traces, separating time spent queueing from time spent working. End with the fix and what stops it happening again.
- Design a banking app: Also reported, and a reminder that the round is not always about models. Here consistency is the point: transactions, idempotent retries, and an audit log you can reconcile.
One candidate's advice, reported by interviewing.io: don't worry about memory constraints, and spend the time on the interesting architectural choices. The prompts rotate, so practise the shapes and say the trade-off out loud each time you make a choice.
What Anthropic screens for
Beyond raw technical skill, Anthropic screens hard for genuine alignment with its safety-first mission, and the values round is the most common place candidates fall short.
- Genuine alignment with AI safety and responsible deployment
- A clear and honest answer to why Anthropic specifically
- Pragmatism and first-principles reasoning over memorized frameworks
- Authenticity and comfort sitting with hard, unresolved trade-offs
- Putting the mission first and acting for the global good
Anthropic interview questions
Candidate-reported themesReported questions split between motivation and ethics on one side and practical, build-from-scratch coding and systems work on the other.
Behavioural & motivation
- Why do you want to work at Anthropic, and why now?Listening for: Specificity · Connection · Research
- How do you think about AI safety, risk, and responsible deployment?Listening for: A view of your own · Sits with trade-offs · Grounded in the work
- Walk through a time you handled ethical friction or disagreement on a team.Listening for: Engagement · Professionalism · Resolution
- Describe a hard technical decision you made and how you weighed the trade-offs.Listening for: The decision and options · First-principles reasoning · Outcome and hindsight
Technical
- Build an in-memory key-value store in Python: set, get and delete first, then range scans, then keys with a time to live. How would you structure it so each level extends cleanly?Build from scratch: A small in-memory database or a key-value file system, in Python, then the data structure underneath it without the libraryListening for: Data model up front · Each level handled · Complexity justified
- Design a concurrent web crawler in Python that fetches, parses and rate limits. Where do concurrency and edge cases bite?Crawlers and parsers: Concurrency and edge cases weigh more than the happy pathListening for: Concurrency model · Edge cases covered · Rate limits and safety
- Serve a large language model efficiently through an API.LLM serving design: An API for serving large models efficiently, request batching, GPU useListening for: Batching and queueing · GPU use under load · Trade-offs said aloud
- Practice: you are shown an LRU cache class where the eviction step removes the most recently used item. Say what the code does, what is wrong with it, and how you would add a size limit in bytes.Reading and extending real code: Say what a snippet does and what is wrong with it, then add to it while defending your complexity choicesListening for: Explains the code · Finds the bug · Extends it soundly
These are the reported themes — your loop is role-specific. Paste the actual posting and Calibrd predicts the questions for that exact role and level.
Predict my Anthropic questions →Pay
Anthropic pays at the top of the market. Levels.fyi reports software engineer total compensation with a median around 746K dollars per year, senior packages near 563K and lead packages near 785K, a large share of it in equity.
How to prepare for an Anthropic interview
- Practise building small working systems end to end in Python, like an in-memory database or crawler, rather than grinding LeetCode, and extend each one twice with a requirement you did not plan for. That is the CodeSignal screen in miniature.
- Prepare the recruiter call as an interview: why this company, in terms of its work and its writing, and what a B Corp is. People with strong CVs fail here.
- Prepare an honest, specific answer for why Anthropic, grounded in its safety work and writing such as Dario Amodei's essays and Anthropic's core views on AI safety.
- Study systems design tied to LLM serving: request batching, GPU memory, KV cache, and multi-region inference.
- Rehearse handling edge cases and concurrency out loud, and be ready to justify your time and space complexity under questioning.
This guide covers Anthropic's engineering and research hiring. For management and leadership roles the loop is similar but the bar shifts to people, delivery and strategy, so pair it with the leadership interview prep hub. The bar for your exact role comes from the role-by-role guides, and the prep that actually transfers is spoken, so run a mock interview before the real one.
Also interviewing at OpenAI, Google DeepMind or Mistral AI? Their guides follow the same format, from the first screen to the offer.
Knowing the questions isn’t the same as answering them out loud. Run an Anthropic mock: spoken answers, coached on the spot. Your first mock is free.
Run an Anthropic mock →FAQ & sources
What is Anthropic's interview process?
Three to four weeks from recruiter call to decision, and faster with a competing offer. The stages, in order: Recruiter screen; CodeSignal assessment; Hiring manager call; Coding, two rounds; System design; Values interview.
What does Anthropic look for in candidates?
Beyond raw technical skill, Anthropic screens hard for genuine alignment with its safety-first mission, and the values round is the most common place candidates fall short. Genuine alignment with AI safety and responsible deployment; A clear and honest answer to why Anthropic specifically; Pragmatism and first-principles reasoning over memorized frameworks; Authenticity and comfort sitting with hard, unresolved trade-offs.
What questions does Anthropic ask in interviews?
Reported questions split between motivation and ethics on one side and practical, build-from-scratch coding and systems work on the other. Technical themes: Build from scratch; Crawlers and parsers; LLM serving design; Reading and extending real code.
What is on the Anthropic CodeSignal assessment?
Anthropic's screen is CodeSignal's project-based format rather than its four-puzzle one: a single system that you build across levels, each level adding requirements on top of the last while every earlier method has to keep working. Ninety minutes, Python with only the standard library, the unit tests for the current level visible from the start, and the next level unlocking only when all of them pass. Covered on this page: In-memory key-value store; Bank with transaction types; File system; Merged entities and historical state; Others reported; Scoring, and why a high one is not an offer of an interview.
What is the Anthropic system design interview?
One hour during the onsite, in a shared Google Doc instead of a drawing tool. The prompts sit close to Anthropic's own work: serving a large model, search over very large corpora, and the pipelines that feed a model. Covered on this page: Design a Claude chat service; Serve a large language model efficiently through an API; Let a model handle several questions in a single thread; A distributed search system over a billion documents at a million queries a second, with LLM inference; Hybrid search that combines keyword retrieval with semantic similarity; An indexing pipeline for a billion documents, with concurrent crawling; p95 latency jumped from 100 ms to 2,000 ms: find it and fix it; Design a banking app.
How do I prepare for an Anthropic interview?
Practise building small working systems end to end in Python, like an in-memory database or crawler, rather than grinding LeetCode, and extend each one twice with a requirement you did not plan for. That is the CodeSignal screen in miniature. Prepare the recruiter call as an interview: why this company, in terms of its work and its writing, and what a B Corp is. People with strong CVs fail here.
- 01Anthropic, raises $65B Series H at $965B valuationthe Series H led by Altimeter, Dragoneer, Greenoaks and Sequoia, May 2026
- 02Anthropic, confidentially submits draft S-1the confidential draft registration statement of 1 June 2026; no share count, price or date set
- 03Anthropic Careersofficial framing on roles, values, and how they interview technical staff.
- 04interviewing.io, Anthropic's interview process and questionsround-by-round loop, the standalone values interview, and the system design round: one hour in a Google Doc, reported prompts (a Claude chat service, multiple questions in one thread, a banking app) and its focus on LLM serving.
- 05Glassdoor, Anthropic interview questionscandidate reports, difficulty rating, and roughly 19-day timeline.
- 06Levels.fyi, Anthropic Software Engineer salarycompensation ranges and medians for engineering levels.
- 07IGotAnOffer, Anthropic interview processsix-step process outline, work-sample coding style, and four reported system design prompts (distributed search at 1B documents, hybrid search, an indexing pipeline, a p95 latency spike); read via the web archive, May 2026.
- 08Sundeep Teki, Anthropic CodeSignal assessment guidethe Industry Coding Framework format, the reported problem types and how each grows by level, the 600-point scale and the 520 line, the integrity review. April 2026.
- 09Sundeep Teki, Anthropic's CodeSignal assessment in 2026three independent candidates in one month reporting six levels rather than four; the rotated problem bank; the advice to spend longer on the data model. June 2026.
- 10Anthropic Engineering, Designing AI-resistant technical evaluationsthe performance engineering take-home in Anthropic's own words, the general rule that take-homes are done without AI unless stated, and the reasons the tests keep being redesigned. January 2026.
- 11r/leetcode, "My Anthropic CodeSignal experience — 580/600 and rejected before interview"the September 2026 account this guide's pacing, preparation method and scoring note come from: four levels on an in-memory banking system, level timings of 11, 17, 47 and the rest, merged entities with historical state as the hardest part, the 95/100 in-test score shown as 580/600 in the portal, the rejection without an interview, and the replies reporting 600 and 595 rejected and 580 advanced.
Interview processes change. This reflects widely-reported and sourced conditions as of 2026 — confirm specifics with your recruiter, and treat it as a map rather than a guarantee.
Prep for a real Anthropic role
Practise your Anthropic interview, out loud.
Paste a real Anthropic posting and Calibrd predicts the questions for that role and level, benchmarks the pay, and flags the gaps an interviewer will probe in your CV — then listens to your spoken answers and coaches them. Your first mock is free.
Free to start · No card · Encrypted at rest, never used to train AI, remove anytime