Microsoft(MSFT)
Search documents
Helios Is AMD’s First AI System To Rival Nvidia Vera Rubin — We Got An Exclusive, First Look
CNBC· 2026-07-20 13:00
Strategic Positioning and Product Launch - AMD launched "Helios," its first rack-scale AI system, featuring 72 GPUs and 18 CPUs per rack to compete directly with Nvidia’s Grace Blackwell and Vera Rubin systems [1][2][3] - The system is designed for high-performance AI inference, emphasizing superior memory bandwidth and customizable configurations compared to proprietary alternatives [4][12] - AMD aims to challenge Nvidia’s 95% dominance in the data center GPU market, targeting a 20% to 25% market share [9][10] Financial Performance and Market Adoption - Data center revenue has become AMD's primary growth driver, increasing 57% year-over-year, with the segment accounting for the majority of total revenue as of Q1 2026 [15][25] - Major tech companies including Microsoft, Meta, Oracle, and OpenAI have signed deployment deals for Helios systems [4][6][9] - Each Helios system is estimated to cost between $5 million and $5.5 million, positioning it as a premium, high-efficiency alternative for AI infrastructure [2][33] Operational and Supply Chain Constraints - AMD is managing significant manufacturing requirements, including the use of TSMC’s advanced 2-nanometer node for MI455 GPUs and securing high-bandwidth memory (HBM) with up to 432GB per GPU [39][40] - To mitigate supply chain risks, AMD committed $10 billion to Taiwanese advanced packaging companies like ASE to secure capacity, following Nvidia's heavy reservation of CoWoS packaging [40] - The company maintains a closed-loop liquid cooling system for Helios, which consumes approximately 750 gallons of chilled water per minute and requires 225,000 to 245,000 watts of power per rack [36][37] Ecosystem and Strategic Development - AMD emphasizes an open-source approach through its ROCm software stack, supporting frameworks like PyTorch, vLLM, and SGLang to provide an alternative to Nvidia’s proprietary CUDA ecosystem [20][29][30] - Strategic acquisitions have been pivotal to the Helios roadmap, including the $50 billion purchase of Xilinx and the $4.9 billion acquisition of server builder ZT Systems [27][28] - Future infrastructure shifts are expected to include the transition from copper cables to optical interconnects within racks over the next 2 to 3 years [41][42]
X @OKX





OKX· 2026-07-20 09:30
Market Outlook - Major technology companies are scheduled to release their financial reports in the upcoming weeks [1] - The list of key technology firms includes Tesla ($TSLA), Alphabet ($GOOGL), Microsoft ($MSFT), Meta ($META), Apple ($AAPL), and Amazon ($AMZN) [2] Strategic Planning - Market participants are advised to maintain continuous monitoring of industry news and corporate updates, as information flow remains active even when financial markets are closed [2]
X @BitMart
BitMart· 2026-07-20 09:26
📊 U.S. Earnings Season Is Here!Big tech companies are about to reveal their latest quarterly results.Earnings surprises, business growth, and future guidance could become key drivers of market moves.🔥 Key Earnings to Watch This Week:🚗 Tesla (TSLAON)📅 Jul 22 After Market Close👀 Focus: Deliveries, margins, FSD progress🔗 Trade Now:https://t.co/yaDOXQxeRg🔍 Google (GOOGLON)📅 Jul 23 After Market Close👀 Focus: AI Search, Google Cloud growth🔗 Trade Now:https://t.co/DuvXhNmcft💻 Microsoft (MSFTON)📅 Jul 24 After Marke ...
Don't Let the LLM Drive - Ornella Bahidika & Joel Allou, Microsoft
AI Engineer· 2026-07-20 06:25
Technical Architecture & Reliability - The core challenge in building multi-step AI agents is not a prompting issue but a control problem, as LLMs struggle to maintain state across complex workflows [1][2] - Implementing a "harness" architecture—a state machine that manages flow, validation, and decision-making—ensures reliability by restricting the model to specific, isolated tasks [2][3][5] - The "harness" acts as the director, while the model serves as the talent, ensuring the model never decides the current state or the next step in a process [2][3][8] Cost & Performance Optimization - Transitioning from heavy frontier models (e.g., Opus 4.7%) to smaller, more efficient models (e.g., Haiku 4.5%) is achievable through harness engineering without sacrificing performance [6] - Leveraging smaller models within a controlled harness significantly reduces latency and operational costs while maintaining the required output quality [4][6] - By providing only the necessary input for a specific scenario, the system minimizes the model's "thinking" overhead, leading to faster and more predictable execution [5][6][10] Strategic Application & Scalability - The harness engineering framework is highly versatile and applicable to various domains, including voice-based AI tutors, coding agents, operational runbooks, and onboarding flows [11] - Adopting this abstraction is recommended for any agent where reliability is currently inconsistent or "a coin flip," shifting control flow away from the model to external logic [11][12] - The design philosophy emphasizes that while models should be utilized for content generation, they should not be permitted to "drive" the overall process or decision-making logic [12]
Your Voice Agent Doesn't Need a Frontier Model - Joel Allou & Ornella Bahidika, Microsoft
AI Engineer· 2026-07-20 06:25
Technical Strategy and Latency Optimization - Real-time voice applications prioritize latency over model intelligence, establishing a strict budget of approximately 950 milliseconds for the AI to begin responding [2] - Frontier models, while superior in reasoning, are often unsuitable for voice interfaces due to multi-second processing delays that disrupt user engagement [2][4][5] - The "Ace" system architecture achieves a response time of approximately 900 milliseconds by utilizing smaller, cost-effective models like Haiku 4.5% instead of larger models like Opus 4.7% [8][9] Architectural Design and Implementation - Logic and reasoning tasks are extracted from the AI model and offloaded to an external state machine, which manages lesson flow, student tracking, and instructional progression [2][3][5][6] - The AI model is restricted to its core competency—natural language generation—by receiving pre-processed summaries and instructions for each turn [3][5][6] - Developers must implement robust "scaffolding" (strict rules and logic) to prevent smaller models from drifting during long-form interactions [10] - The cost of building external scaffolding is a one-time development investment, which ensures system stability without incurring additional latency per interaction [10] Industry Insights and Best Practices - For high-volume, real-time applications, the optimal strategy is to select the fastest model that fits the latency budget and dedicate development resources to external logic harnesses [11][12] - The AI model should be treated as the smallest component within a larger, highly structured system to ensure performance and scalability [12][13]
Build the AI GTM Agent That Knows the Buyer - Dr. Sajjan Kanukolanu, Position2 (Position Squared)
AI Engineer· 2026-07-20 06:25
Hello. We know that the modern B2B selling has evolved significantly. The uncomfortable truth is that by the time the buyer reaches you, the decision is mostly made.They've researched, compared, and probably shortlisted vendors or partners that they want to work with. That may include you or your competitors. This talk is about the architecture that lets GTM teams know the buyer more than they today.My name is Sajjan Khan Akolanu. I have 20-plus years experience across product, technology, and marketing. At ...