What you need to know

  • Anthropic confirmed an in-house silicon team on 5 August — the first public acknowledgement of a chip effort the company had previously kept quiet.
  • The pitch is co-design: silicon and Claude models developed together, with reports citing a target of roughly 50% lower per-token inference costs.
  • Nothing exists yet. This is a hiring announcement, not a product. Manufacturing talks with Samsung — reportedly around its 2 nm SF2P node — remain exploratory.
  • The multi-vendor stack stays. Claude will keep running on AWS Trainium, Google TPUs, NVIDIA GPUs and AMD hardware; Anthropic's own part would slot in beside them.
  • The salary band tells its own story: chip roles advertised at $320,000 to $485,000, and the premium leaks into every silicon-adjacent software skill.

The confirmation, and what took so long

On 5 August 2026, Anthropic published job listings that did something it had never done before: publicly admit it is designing its own chips. "We work from the chip level up with our silicon partners, and we are now deepening that investment by building a custom silicon team," the listing reads. Reuters had reported that the company was exploring custom silicon back in April, per Unite.AI's account, but this is the first time Anthropic has said it out loud.

The roles span the full silicon lifecycle: front-end design, pre-silicon verification, physical design, design-for-test, analogue and mixed-signal, foundry technology, design infrastructure, and packaging with signal and power integrity. That last cluster matters — packaging is where modern accelerator projects live or die, and hiring for it signals a team that intends to ship hardware rather than publish research. The listings ask for engineers who have personally contributed to shipped silicon, and pay accordingly: $320,000 to $485,000.

The team also reportedly has a notable hire at its core: Clive Chan, described in several reports as one of the early engineers on OpenAI's custom silicon programme, is said to have joined Anthropic to help build the effort. Anthropic has not commented on individual hires, so treat that as reported rather than confirmed.

Why co-design is the real story

It would be easy to file this under "another lab tries to escape NVIDIA margins". That is part of it, but the more interesting claim is architectural. Anthropic is not buying a merchant chip and porting Claude to it; it says it will design the silicon and the model together. When the same organisation controls both sides, it can make trades that are impossible at arm's length: quantisation formats the model is actually trained for, memory hierarchies sized to the KV-cache behaviour of a specific architecture, interconnect topologies matched to how the model is sharded in production.

That is where the reported target of roughly 50% lower per-token inference costs comes from — a figure carried by TechTimes and others. It is a target, not a benchmark; no silicon exists to measure. But it is not a fantasy number either. Google has run its frontier models on co-designed TPUs for a decade, and OpenAI's Jalapeño programme was built on the same logic. General-purpose GPUs spend die area and power on flexibility that a single-customer inference chip simply does not need.

The economics behind the bet are straightforward. Anthropic's revenue run-rate reportedly passed $30 billion annualised this spring, and inference is the dominant marginal cost of serving that demand. Halve the per-token cost on even a fraction of that workload and the chip pays for its own development inside a generation — even at leading-edge design costs, which one industry estimate cited by FourWeekMBA puts at $500–750 million for a 2 nm or 3 nm part.

Pro tip

Watch what co-design does to which tokens get cheaper. Custom inference silicon tends to favour high-batch, steady-state serving — the economics of cached, long-context and batch API traffic improve first. If your workload fits batch processing, you are positioned to inherit these savings early; latency-critical single-stream traffic usually benefits last.

How the in-house silicon bets compare

Every frontier lab now has a position on custom chips. The instructive differences are in scope and posture.

Lab Programme Status (Aug 2026) Posture
Anthropic Unnamed inference chip, co-designed with Claude Design team confirmed 5 Aug; Samsung manufacturing talks exploratory Additive — one lane in a multi-vendor stack
OpenAI Jalapeño, with Broadcom Unveiled 24 June; deployment from late 2026, 10 GW buildout to 2029 Purpose-built inference fleet at scale
Google TPU family Seventh generation; in production internally since 2015 The template — a decade of model-chip co-design
Amazon Trainium / Inferentia (Annapurna Labs) Deployed at scale, including large Anthropic clusters Cloud-vendor economics, sold as a service
Meta MTIA In production for ranking and recommendation workloads Internal workloads first, LLM serving later

The sharpest contrast is with OpenAI. Jalapeño, unveiled with Broadcom on 24 June, is a reticle-sized inference-only ASIC that went from design to production in around nine months, per Tom's Hardware — with deployment beginning late this year. OpenAI built a single big part with a single partner and committed a 10-gigawatt buildout to it. Anthropic is explicit that its silicon will sit inside a multi-vendor stack rather than replace it. One is a conviction bet; the other is a hedge with upside.

It also lands in a market that is busily unbundling inference from general-purpose GPUs from every direction — from Etched's transformer-only ASIC to OLIX's photonic inference part that skips HBM entirely. The labs and the startups have converged on the same conclusion: inference is where the money burns, so inference is where the silicon goes.

Samsung, 2 nm and the manufacturing question

Designing a chip and getting it fabricated are different problems, and the second is where Anthropic's programme is least settled. Reports going back to early July — first from TechCrunch, then TechTimes and others — describe exploratory talks with Samsung about manufacturing on its 2 nm process, specifically the performance-tuned SF2P node with gate-all-around transistors. Those reports also describe the talks as genuinely early: what the chip does, how it fits in a server and how powerful it needs to be were all reportedly still open questions.

A Samsung deal would be notable in its own right. TSMC fabricates nearly every serious AI accelerator today, and its leading-edge capacity is heavily committed. A frontier lab validating Samsung's 2 nm foundry line would diversify a supply chain the whole industry is nervously dependent on — one reason several observers expect Anthropic to try to firm up the relationship before the year is out.

Watch out

Do not price unannounced silicon into your 2027 unit economics. Between team formation and volume deployment sit tape-out, bring-up, yield, packaging, and the software stack — each a place where schedules slip by quarters. Budget on today's Claude pricing and treat any chip-driven cut as upside, not baseline.

The multi-vendor stack is not going anywhere

Anthropic currently runs Claude across four external compute lanes: AWS Trainium, Google TPUs, NVIDIA GPUs and AMD accelerators. The company has been explicit that its own chip would be a fifth lane, not a replacement — a posture that makes sense when your constraint is total capacity rather than any single vendor's pricing. Google alone is reported to be bringing well over a gigawatt of additional TPU capacity to Anthropic from 2027.

For everyone downstream, the direction of travel matters more than the timeline. Inference capacity is the choke point of this cycle — it is why GPU lease prices have been rising even as model prices fall, and why lenders have started treating inference chips as collateral, as we covered when the first inference-chip-backed loan closed. Every credible new lane of non-NVIDIA inference supply — lab-built or startup-built — puts downward pressure on the price of serving tokens, whoever's logo is on the die.

What cheaper Claude tokens would mean in India and the UK

Per-token cost is the single number that decides which AI products are viable outside Silicon Valley pricing tolerance. A Bengaluru or Pune startup serving Indian consumers at Indian price points, and a London fintech running compliance summarisation inside regulated margins, hit the same wall: gross margin per request. Halve the provider's cost of inference and some of that saving historically reaches the API price — the past two years of falling frontier-model prices suggest competition passes a meaningful share through.

Concretely, a durable cost reduction on Claude serving would move three things for builders in both markets. First, agentic workloads — the token-hungriest category — become viable at consumer price points, not just enterprise ones. Second, always-on background inference (monitoring, enrichment, evaluation loops) stops being a luxury line item. Third, the calculus of self-hosting open-weight models shifts again: every API price cut raises the bar that a self-managed GPU deployment has to clear, which matters in India where teams often self-host to control costs, and in the UK where teams often self-host for data-residency reasons.

None of that arrives this year. But builders making 2027 platform bets should register the direction: every frontier lab is now engineering the cost of tokens downwards at the silicon level, and pricing roadmaps will follow.

Every article here is written by a Verified Builder. Want your name on the next one?

AI Tech Connect lists AI engineers, founders and researchers across India and the UK — and the people hiring browse it to find them. Adding your profile is free.

Become a Verified Builder →

The $485,000 signal: silicon-adjacent careers

The most immediately useful datapoint in this story for most readers is the salary band. Anthropic is advertising chip roles at up to $485,000 — and it is bidding against OpenAI, Google, Amazon, Meta, NVIDIA and a wave of funded ASIC startups for a talent pool that barely grew during the decade when software ate the industry's attention. Scarcity plus five simultaneous frontier-lab silicon programmes equals a repricing, and it does not stop at chip designers.

If you are a hardware engineer

Physical design, verification, packaging and signal-integrity skills are now frontier-lab skills. India's semiconductor design workforce is one of the world's largest — much of it doing back-end work for the established vendors from Bengaluru and Hyderabad — and the UK retains deep accelerator design talent around Cambridge and Bristol. The labs hiring at these bands recruit globally, and remote-plus-relocation offers follow scarcity. A public track record of shipped silicon is now among the highest-value credentials in AI.

If you write software

Co-design programmes need far more than RTL engineers. Kernel authors, compiler engineers, inference-runtime developers and people who can reason about serving economics end to end are the connective tissue between model teams and chip teams. The ecosystem-level demand for portability layers is part of why Qualcomm paid $3.9 billion for Modular's CUDA alternative — the industry is paying a premium for anyone who can make models fast on hardware that is not an NVIDIA GPU. If you are choosing what to learn next quarter, performance engineering close to the metal is being repriced upwards in both India and the UK.

What to watch

Three markers will tell you whether this programme is real progress or an expensive hedge: whether the Samsung relationship firms up into a signed manufacturing deal by early 2027; whether Anthropic's silicon hiring accelerates beyond the founding team; and whether the company starts talking about a target workload — a co-designed chip only halves costs on the traffic it is designed for. Until then, the safe read is the one Anthropic itself is making: the multi-vendor stack pays the bills, and the chip team is a bet on what comes after.