← Back to glossary
+Suggest a term
Tool·Infrastructure·Added today

Etched

Also known as: Etched AI, Sohu chip, Etched Sohu

An AI hardware startup building Sohu, a transformer-specific ASIC (application-specific integrated circuit) designed exclusively for transformer inference. The bet: hardwiring transformer attention into fixed silicon is dramatically faster and cheaper than running the same workload on general-purpose GPUs.

Most AI inference today runs on NVIDIA GPUs, which are flexible general-purpose accelerators that can run any computation. Etched's Sohu chip takes the opposite approach: it implements the transformer architecture directly in silicon, with no support for other workloads. Because every transistor is dedicated to one job, the company claims Sohu can run transformer inference roughly 20 times faster than an H100 GPU at significantly lower energy cost, though as of mid-2026 these figures are vendor claims without independent verification.

Etched emerged from stealth on June 30, 2026, announcing roughly $800 million raised, working silicon, and over $1 billion in signed customer contracts before shipping a single production rack. First racks were scheduled to ship in summer 2026. The company is targeting AI inference providers and hyperscalers that run transformer-heavy production workloads where the performance advantage outweighs the loss of flexibility versus a general-purpose GPU.

The key risk is also the key premise: Sohu only makes sense if transformers remain the dominant AI architecture for long enough to recoup the specialization investment. If a meaningfully different architecture displaces transformers, Sohu's fixed function becomes a liability rather than an advantage. Builders evaluating inference infrastructure should understand Etched as a high-conviction architectural bet, not a general-purpose compute option.

This definition is AI-generated and refreshed weekly. It may contain inaccuracies. Use your own judgment, especially for production decisions.
Related terms
LPUGPU / TPUInferenceServerless inferenceModel serving