Together AI interview, decoded.

Updated · Sources at the end

Together AI runs a cloud for open-source AI models: inference that customers call through an API, fine-tuning and reinforcement learning, and GPU clusters for training. Its engineering interviews follow from that. Candidates report ordinary coding rounds, then a round that is specific to this company: writing working model-serving code, such as an attention step, request batching or streaming generation, and explaining the performance trade-offs. Inference system design and conversations about past projects follow.

The company raised $800 million in July 2026 at an $8.3 billion valuation, and counts Cursor, Cognition and Decagon among its customers. It is hiring across San Francisco, Amsterdam, London, Bangalore and Singapore.

Free · 2 minutes · no account

The questions for your exact Together AI role and level.

See the questions Together AI asks ↓Run a free mock interview →

Interviewing at Together AI? Below are the questions candidates report. For the ones your exact role and level will get, paste the posting. Your first mock is free.

01

Who Together AI hires

Most open roles are in engineering: inference and compute infrastructure, backend platform, distributed ML systems, networking and AI infrastructure systems, plus research engineers who work on post-training and inference engines. The postings ask for strong software engineering in Python, Go, Rust or similar, and for inference roles hands-on experience with engines such as SGLang, vLLM and TensorRT-LLM. This guide follows the engineering path; research roles add research rounds.

02

What is the Together AI interview process?

Together AI does not publish its interview process. The shape below comes from candidate reports collected by interview-prep sites, and every round is marked as reported. The sources agree on the coding rounds and the applied ML systems round; they differ on whether a take-home comes before the onsite. Reports put the whole loop at about two to four weeks.

01
Recruiter call
About 30 minutes, video

Your background, why Together AI, and which team. Candidates report questions about your experience with models, GPUs or serving systems (reported).

02
Technical screen
About 60 minutes, live coding

One or two coding problems at medium difficulty, in Python, C++ or CUDA depending on the role. One account describes two medium problems in about 45 minutes (reported).

03
Take-home, some roles
Reported at 4 to 8 hours

A realistic engineering problem. Only one prep site describes this step, so treat it as possible, not certain (reported).

04
Coding interview
About 60 minutes

An algorithms round on arrays, strings, hash maps and intervals, solved at speed (reported).

05
Applied ML systems round
About 60 minutes, you write code

Practical model-serving code: implementing an attention step, handling streaming token generation, batching requests efficiently, and explaining the performance trade-offs. Inference-engine roles can get a CUDA kernel to write or optimise (reported).

06
System design
About 60 minutes

Inference or training infrastructure: serving many open-source models on shared GPUs, multi-tenant fine-tuning, KV-cache and batching choices (reported).

07
Team and hiring manager interviews
Two to four conversations

A detailed review of one or two past projects: what you built, what failed and what you measured, plus how you work with other teams (reported).

03

Together AI system design: inference infrastructure

Reported prompts and how to take them

The system design round is about the platform the company runs. Prep sites describe prompts rather than exact questions, so practise the shape: many models, shared GPUs, many customers, and a latency target.

  • Serve 100+ open-source models on shared GPU capacity: Start from the traffic: which models are hot, their sizes and latency targets. Cover placement (which models share a GPU), autoscaling on queue depth, cold-start cost of loading weights, and routing. Name what you would measure first.
  • A multi-tenant fine-tuning service with LoRA adapters: Keep one base model in memory and swap small adapters per request; discuss how many adapters fit, how to batch requests across adapters, and how to isolate customers' data.
  • A KV cache for many concurrent chats: Size the cache per token and per request, then choose eviction and paging (as in paged attention), and say what happens to latency when memory runs out.

These prompts come from interview-prep sites, not from Together AI. Use them to practise the shape of the round.

04

What Together AI screens for

The postings and the candidate reports point to the same few things.

  • Working code over theory: the ML systems round asks you to write it, not describe it
  • Performance judgement: explaining latency, throughput and memory trade-offs in what you built
  • Hands-on inference experience with engines such as SGLang, vLLM or TensorRT-LLM, which the inference postings ask for
  • Ownership of production systems, including on-call for a platform that runs around the clock
  • Measured results in past projects, with what failed as well as what worked
05

Together AI interview questions

Candidate-reported themes

No candidate account we could read lists exact questions, and prep sites describe prompts rather than verbatim problems. What follows is built from the rounds candidates describe, so what you practise matches what they score.

Behavioural & motivation

  • Walk me through a system you built that serves traffic in production. What did you measure, and what broke?Listening forA real system · What they measured · What failed
  • Tell me about a time you made a trade-off between latency and cost.Listening forThe trade-off · The decision · The result
  • Why Together AI, and why this team?Listening forSpecificity · Connection · Research

Technical

  • Practice: implement scaled dot-product attention for a batch of sequences in PyTorch or NumPy, with a causal mask, then say how you would make it faster.Attention from scratchReported in the applied ML systems roundListening forCorrect · Its cost · Faster
  • Practice: requests with different prompt lengths arrive at an inference server. Write a scheduler that batches them to keep the GPU busy without hurting latency.Batching requestsReported in the applied ML systems roundListening forA batching policy · Limits · Fairness
  • Practice: stream tokens to a client as the model generates them, and handle a client that disconnects mid-stream.Streaming generationReported in the applied ML systems roundListening forStreams correctly · Handles cancel · Backpressure
  • Practice: given a list of time intervals, merge the overlapping ones and return the result sorted.Algorithms at speedReported in the coding rounds: two medium problems in about 45 minutesListening forApproach · Complexity · Edge cases

These are the reported themes — your loop is role-specific. Paste the actual posting and Calibrd predicts the questions for that exact role and level.

Predict my Together AI questions →
06

Pay

Software engineer, inference / compute · posted US base$160–280K
Research engineer, post-training inference · posted US base$200–290K
Engineering and research interns · reported$58–63 / hour

Together AI posts a US base range on its job ads, with equity and benefits on top (read 29 September 2026). The software engineer range covers junior to staff, so where you land depends on level. Its careers page lists salary and equity packages, 401(k) matching, meals and a relocation stipend for moves to the Bay Area. Calibrd benchmarks an offer against your level and location in the report.

07

How to prepare for a Together AI interview

  1. Write attention, a batching loop and a streaming endpoint from scratch before the loop. Candidates report writing this kind of code, not only discussing it.
  2. Be ready to explain every performance choice in numbers: latency, throughput and memory.
  3. If you are applying for an inference role, know one engine well (SGLang, vLLM or TensorRT-LLM), including how it batches and caches.
  4. Pick one or two production projects and prepare them in depth: scale, what you measured, what failed.
  5. Practise medium algorithm problems against the clock: two in about 45 minutes is the pace reported.
  6. Ask the recruiter whether your loop includes a take-home; reports disagree.

This guide covers Together AI's engineering and research hiring. For management and leadership roles the loop is similar but the bar shifts to people, delivery and strategy, so pair it with the leadership interview prep hub. The bar for your exact role comes from the role-by-role guides, and the prep that actually transfers is spoken, so run a mock interview before the real one.

Also interviewing at Crusoe, Harvey or Cursor? Their guides follow the same format, from the first screen to the offer.

Knowing the questions isn’t the same as answering them out loud. Run a Together AI mock: spoken answers, coached on the spot. Your first mock is free.

Run a Together AI mock →
08

FAQ & sources

What is Together AI's interview process?

Together AI does not publish its interview process. The stages, in order: Recruiter call; Technical screen; Take-home, some roles; Coding interview; Applied ML systems round; System design; Team and hiring manager interviews.

What does Together AI look for in candidates?

The postings and the candidate reports point to the same few things. Working code over theory; Performance judgement; Hands-on inference experience with engines such as SGLang, vLLM or TensorRT-LLM, which the inference postings ask for; Ownership of production systems, including on-call for a platform that runs around the clock.

What questions does Together AI ask in interviews?

No candidate account we could read lists exact questions, and prep sites describe prompts rather than verbatim problems. Technical themes: Attention from scratch; Batching requests; Streaming generation; Algorithms at speed.

What does Together AI ask in system design?

The system design round is about the platform the company runs. Prep sites describe prompts rather than exact questions, so practise the shape: many models, shared GPUs, many customers, and a latency target. Covered on this page: Serve 100+ open-source models on shared GPU capacity; A multi-tenant fine-tuning service with LoRA adapters; A KV cache for many concurrent chats.

How do I prepare for a Together AI interview?

Write attention, a batching loop and a streaming endpoint from scratch before the loop. Candidates report writing this kind of code, not only discussing it. Be ready to explain every performance choice in numbers: latency, throughput and memory.

Prep for a real Together AI role

Practise your Together AI interview, out loud.

Paste a real Together AI posting and Calibrd predicts the questions for that role and level, benchmarks the pay, and flags the gaps an interviewer will probe in your CV — then listens to your spoken answers and coaches them. Your first mock is free.

Free to start · No card · Encrypted at rest, never used to train AI, remove anytime

Two weeks out?

Email me the Together AI prep checklist

The round that decides it, the questions they ask, and what to do before, in one email. Nothing else unless you sign up.

Together AI Interview: ML Systems Round, Loop, Pay — Calibrd