AMD Inference Microservices
Search documents
AMD AI Workbench - AMD Inference Microservices (AIMs)
AMD· 2026-07-24 12:34
Technology Features & Capabilities - AMD AI Workbench introduced AIMs (AMD Inference Microservices) to simplify AI model serving for automated, scalable, and optimized inference [1] - The platform features an AIM catalog containing a curated and regularly updated collection of open-source models from the Hugging Face community, exemplified by the deployment of the Mistral 24B model [2] - Users can optimize deployment performance by configuring default metrics or selecting custom metrics, alongside enabling auto-scaling options such as replica ranges [3] Deployment & Management Workflow - Model deployment supports public models without requiring tokens, while gated models necessitate a Hugging Face token which can be selected or newly added [3][4] - The Deployed Models tab and Dashboard page allow users to track download progress, oversee project workloads, monitor GPU usage, and check deployment status [5] - Users can interact with deployed models through a built-in chat interface, which also supports comparing multiple models simultaneously [5][6] Integration & API Utilization - AIMs provide an OpenAI-compatible API for large language models to facilitate seamless application integration [7] - Connection details, including external and internal URLs alongside sample Python code, are accessible via the model connection dialog [8] - API keys can be generated, assigned to specific models, and integrated into external code environments to enable remote model prompting and inference [9][10]