AMD(AMD)
icon
Search documents
An Open Foundation for the Age of AI-Powered Robots
AMD· 2026-08-14 12:23
Market Dynamics and Industry Trends - AMD positions itself as an open foundation for the age of AI-powered robots [1] - AMD maintains active engagement across major digital and social media platforms including Discord, Facebook, Twitter, Twitch, LinkedIn, and Instagram [1] Intellectual Property and Corporate Identity - Advanced Micro Devices, Inc. holds trademark rights for AMD, the AMD Arrow Logo, and related combinations in the United States and other jurisdictions as of 2026 [1]
Training at Scale with AMD Primus
AMD· 2026-08-14 12:23
Technology Capabilities - Primus enables reliable, debuggable, and high-performance large-scale training on AMD Instinct GPUs [1] - The platform supports the latest open-source training frameworks, models, and is expanding compatibility to cutting-edge model architectures, training techniques, and data types [1] Market Competitiveness - Primus delivers state-of-the-art pre-training and post-training performance proven at scales of thousands of GPUs [1] - Advanced Micro Devices positions the AMD Instinct GPU as a competitive solution for model development at frontier laboratories, enterprises, and artificial intelligence startups [1]
Inside AMD Helios: Architecture of a Rack-Scale AI System
AMD· 2026-08-14 12:23
Infrastructure Strategy - AMD introduced Helios as a rack-scale AI infrastructure designed to accelerate training and inference for next-generation AI workloads [1] - The architecture encompasses compute, networking, memory, and system design considerations [1] Market & Economic Impact - Rack-scale optimization enhances performance, scalability, efficiency, and total cost of ownership (TCO) for enterprise and cloud AI deployments [1]
Evolving Comms Libraries in ROCm for Future AI Workloads
AMD· 2026-08-14 12:22
Application Trends and Challenges - AI training models have evolved from simple single-GPU setups to complex architectures requiring tensor parallelism, pipeline parallelism, and mixture of experts (MoE) [1][3][4] - Inference applications demand low latency where every microsecond impacts user experience and costs, requiring the elimination of staging buffers and direct GPU-to-GPU High Bandwidth Memory (HBM) writes [5][6][16] - Large-scale training running across tens of thousands of GPUs faces critical challenges including network fabric congestion, job bootstrap time increasing to tens of seconds or minutes, and memory resource consumption on High Bandwidth Memory (HBM) and qubits [7][8][9] Rekall Library Innovations - AMD collective communication library Rekall utilizes system Direct Memory Access (SDMA) copy engines to move data across GPUs, freeing up compute units for matrix-matrix multiplication and achieving significant speedups [10][11][15] - Single-node optimizations for small messages implement one-shot and two-shot algorithms, achieving significant speedups across different collectives with a maximum speedup of 3.7% on an 8-GPU MI350 node [17][19][20] - Multi-node communication implements a topology-aware hierarchical algorithm that achieves a speedup of up to 3.6% (or 3.6 GB) across 16 nodes with 8 GPUs (MI350) per node [21][23][24] - GPU-initiated communication enables compute and communication fusion, using a single Compute Unit (CU) via SDMA to achieve higher or similar performance compared to default algorithms using up to 64 CUs [25][26][28][29] - Congestion-aware spray traffic (CAS) technology measures round-trip times across queue pairs to dynamically distribute load, achieving bandwidth close to optimal and a speedup of up to 1.7% in multi-node all-to-all operations [35][36][37][39] Roxamen Runtime Performance - Roxamen acts as a GPU-initiated runtime implementing OpenCMN APIs, utilizing SDMA with a single Compute Unit (CU) to deliver up to 35 times higher bandwidth for larger messages on 8 MI355 GPUs compared to default algorithms using 64 CUs [41][42][43] - Supports multiple Network Interface Cards (NICs)—including AMD Pensando Polara, Thor, and Connect7—achieving high bus bandwidth close to 50 gigabytes per second and highly competitive latencies for 8-byte messages [44][45][46]
Accelerating LLM Inference on AMD ROCm with AITER and ATOM
AMD· 2026-08-14 12:22
For the next about 20 minutes, I'll talk about our inferencing framework called ATOM. And what's behind that is a collection of the kernels, which is called AITER. I think earlier today, when you hear about the keynote speech, you might realize that AMD has delivered a lot of the very good hardware.But to light up those GPUs, we actually need a very mature and sophisticated software stacks to make sure that we could leverage whatever that they offer by our hardware. So that's exactly what I'm going to prese ...
Transforming Ethernet from Underdog to Champion for AI Inference
AMD· 2026-08-14 12:22
Modern AI inference is increasingly constrained by KV cache movement across prefill, decode, and storage tiers rather than GPU compute. This session demonstrates how a lightweight software layer enables RDMA-class performance on standard Ethernet networks without application changes, supporting high-performance disaggregated vLLM inference. Live benchmarks highlight improvements in time-to-first-token (TTFT) and inter-token latency (ITL). Discover more: https://www.amd.com/en/corporate/events/advancing-ai/s ...
vLLM in 2026: Challenges and Optimizations
AMD· 2026-08-14 12:21
Thank you so much. Thank you for showing up and coming with me to join this session. And what we plan to do today, over the next 20 minutes, is to really cover.there's a lot to cover about inference, but I would like to sort of tell you a little bit about how we approach open source inference from vLLM, which is the inference engine that lets you run any large language model on data center hardware. And really there to. hopefully this will be a start and a spark of discussion and inspiration for you all.And ...
OpenJarvis: Personal AI, on Personal Devices
AMD· 2026-08-14 12:21
Market Trends & Industry Dynamics - Personal artificial intelligence is becoming central to daily work, though most current systems continue to operate within cloud infrastructure [1] - Local open-weight models trail frontier cloud models by **6 to 12 months** in capabilities, while consumer hardware accelerators currently support open-weight models ranging from **1 to 128 billion** parameters [4] Investment Opportunities & Cost Efficiency - Operating personal artificial intelligence on-device substantially reduces daily expenditure, achieving up to an **800x** reduction in financial cost and significantly lowering latency compared to cloud-only stacks [21] - Collaborating with cloud resources allows local models ranging from **20 to 30 billion** active parameters to rival closed-source frontier models while reducing daily operational costs by **7x to 11x** [23][24]
Efficient General Intelligence - When Novel Models Meet Customized Silicons
AMD· 2026-08-14 12:21
All right, good morning. Glad to be here. It's great to see that this AMD event is so well attended.So I'll be talking about a topic a lot of us are very, very interested in, on Efficient General Intelligence. I want to emphasize on the efficiency part. The reason is that there's a lot of exciting progress on kind of how much of an intelligence we can get with the latest frontier models we have.However, we have to say that if you look at the efficiency, there's a huge gap. For example, if you take some rece ...
Building Next-Gen AI Infrastructure: Scaling Enterprise LLM Serving with RadixArk
AMD· 2026-08-14 12:21
Thank you, everyone, for joining this session. I'm very excited to talk about the SGLN Miles, two open source frameworks that we have spent time building on this that aim to help you build frontier AI infrastructure. Let's start with, I mean, what is SGLN.Some of you might already heard of and maybe are using it. It's an open source framework for inference and has been widely adopted in production. People use it to serve open source models in many companies across all the categories.It aims for a target for ...