Amazon Announces Its Own AI Chips For Growing Demand
Amazon is advancing its investment in AI chips to address the increasing demand for efficient computing. The company's Inferentia chips are central to this initiative, powering Amazon EC2 Inf1 and Inf2 machines designed for running deep learning and…
Amazon is advancing its investment in AI chips to address the increasing demand for efficient computing. The company's Inferentia chips are central to this initiative, powering Amazon EC2 Inf1 and Inf2 machines designed for running deep learning and generative AI at scale.
The first generation of Inferentia chips offers up to 2.3 times higher throughput and 70% lower cost per inference compared to similar EC2 machines. Companies such as Finch AI, Sprinklr, Money Forward, and Amazon Alexa have utilized these machines to reduce operational costs.
The updated Inferentia2 chips enhance performance significantly, offering up to four times higher throughput and ten times lower latency than the previous model. These chips also include enhanced memory support, providing 32GB of HBM per chip, which is four times the capacity of the earlier version. This allows for the execution of more extensive models, ranging from language systems to image generators.
Developer Support with Inferentia Chips
Amazon integrates the hardware with its Neuron software kit, enabling developers to run models without complete redevelopment. Neuron is compatible with PyTorch and TensorFlow, allowing teams to maintain their existing workflows.
The tool automatically converts high precision FP32 models into lower precision formats such as FP16, BF16, INT8, and the new FP8 option available with Inferentia2. This process reduces the time required to bring models into production since manual model retraining is unnecessary. The chip also supports dynamic input sizes, custom C++ operators, and stochastic rounding to enhance accuracy during intensive workloads.
Amazon is advancing its investment in AI chips to address the increasing demand for efficient computing.
The range of tasks that these chips can handle is extensive. Customers use Inferentia for applications including language processing, speech recognition, image generation, and fraud detection.
In addition to Inferentia, Amazon is enhancing training capabilities with its Trainium3 UltraServers. Announced at the re:Invent event, the new system incorporates up to 144 Trainium3 chips based on a 3nm process, delivering up to 4.4 times more compute performance and four times better energy efficiency than Trainium2 UltraServers.
Clients such as Anthropic, Karakuri, Metagenomi, NetoAI, Ricoh, and Splash Music have employed Trainium chips to reduce training and inference costs by up to 50%. Decart has achieved real-time generative video processing four times faster and at half the cost of GPU machines. Amazon Bedrock is already utilizing Trainium3 for live workloads.
Trainium3 provides three times higher throughput per chip and four times faster response times. These improvements reduce training durations from months to weeks, accelerating the delivery of new AI products to customers.
Additional AI Announcements from Amazon
As part of a broader initiative unveiled at re:Invent, Amazon introduced new Nova models and launched Nova Forge, providing customers with access to model checkpoints for data integration. Companies like Reddit and Hertz are leveraging these tools to expedite automation and development processes.
Amazon also introduced frontier agents capable of operating for extended periods without user input. These agents are designed for software development, security, and DevOps tasks. Early adopters, including Commonwealth Bank of Australia and SmugMug, have utilized them to enhance team productivity.
Furthermore, Amazon announced AWS AI Factories to integrate Trainium chips, NVIDIA GPUs, and Bedrock services directly into customer data centers. In Saudi Arabia, HUMAIN plans to establish an "AI Zone" with up to 150,000 AI chips using this configuration, underscoring Amazon's commitment to expanding its hardware reach.
Based on reporting by techround.co.uk.
