Lifestyle

AWS introduces Bedrock AgentCore Runtime Instances: AI agents can run for 14 days, use GPUs and collaborate on one machine

In a 2026-09-30 post, AWS used a music production pipeline to demonstrate Runtime Instances, a new way to host AI agents on Amazon Bedrock AgentCore. It offers sessions of up to 14 days, GPU access, lasting storage and several agents working together on one machine. It matters mainly to developers and businesses building AI agents. All information comes from a single source: AWS.

About 7 min read

AWS introduces Bedrock AgentCore Runtime Instances: AI agents can run for 14 days, use GPUs and collaborate on one machine
Image: Mokaair (Original editorial artwork)

What happened

On 2026-09-30, AWS published a post on its Artificial Intelligence blog explaining how to build a multi-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances. An AI agent is software that uses an AI model to carry out a task step by step. According to AWS, AgentCore now offers two compute options for hosting AI agents. The first is serverless MicroVMs, where AWS runs small, short-lived machines for you. The second is Runtime Instances, which AWS calls the "new option". Runtime Instances use AWS-managed EC2 infrastructure (EC2 is Amazon's rentable cloud servers) and are aimed at persistent, long-running agent workflows.

AWS notes that as organizations move from single-purpose agents to multi-agent systems, infrastructure needs change. If multiple agents must share context across work spanning several days, serverless sessions capped at a few hours are not enough. The post does not give an official launch date for Runtime Instances, so the event date here is the post's publication date.

AWS introduces Bedrock AgentCore Runtime Instances: AI agents can run for 14 days, use GPUs and collaborate on one machine
Mokaair editorial verification flow · Image: Mokaair (Original editorial artwork)
Read the full description

Sources are collected, independently checked, then reviewed by Jev.

How MicroVMs and Runtime Instances differ

AWS says both options support agent-building frameworks such as CrewAI, LangGraph, LlamaIndex and Strands Agents. Both also let you choose your foundation model, the underlying AI model, and integrate with MCP and A2A, which are protocols for connecting agents to tools and to each other. The main difference lies in the underlying compute model. The table below summarizes the comparison in the AWS post.

Comparison of AgentCore's two compute options (source: AWS blog). A GPU is the chip used to run AI models; Savings Plans and ODCRs are AWS discount and reserved-capacity arrangements.
ItemMicroVM (serverless)Runtime Instances
ComputeFully managed by AWSAWS-managed EC2 instances
Session lengthUp to 8 hoursUp to 14 days
Agents per compute unitOne agent per microVM (1:1)Multiple agents per EC2 instance (1:N)
Artifact typesContainer images and Amazon S3 sourcesContainer images and Amazon S3 sources
GPUNot supportedSupported (on supported instance families only)
PersistenceSession-scopedAmazon EBS persistent storage
PricingPay per useEC2 runs in your account; Savings Plans and ODCRs can be used
ScalingOn demandManaged by a capacity provider

According to AWS, the key mechanism is the "shared session". Each agent is deployed as its own runtime. A capacity provider is the setting that tells AgentCore which machines to set up. When two agent runtimes use the same capacity provider and are invoked with the same runtimeSessionId (a session identifier), AgentCore places them on the same EC2 instance. There they share a file system and can read each other's output directly.

AWS's demo: three agents collaborate on a track

  • Composer agent: according to AWS, it uses Claude Sonnet 4.6 to turn the producer's request into a music brief. It then generates audio on the instance's GPU with ACE-Step, an open-source music generation foundation model.
  • Delivery agent: reads and measures the track on the shared file system. It then asks Claude Sonnet 4.6 to plan EQ, compression and limiting (standard audio adjustments to tone and volume) based on the measurements. It applies them and measures again to confirm the targets are met.
  • Compliance agent: independently re-measures the finished track and checks the delivery targets. It also compares harmonic similarity against the studio's own catalog. If it finds a similarity, it calls back the composer agent to produce an alternative version.

AWS published sample results on a g6.xlarge instance in the us-east-2 region. Preparing the model stack took 239 seconds. Composing took 25 seconds, including 8.98 seconds of rendering on an NVIDIA L4 with peak VRAM (GPU memory) of 7.63 GiB. Delivery took 41 seconds and compliance screening took 28 seconds, with all five steps completed on the same instance. Before delivery processing the audio measured -7.5 LUFS (a loudness measure) with a 0.42 dBTP peak (the highest signal level). Afterward it measured -14.0 LUFS with a -3.2 dBTP peak. The compliance screening result was "REVIEW REQUIRED" (human review needed). These figures come from a single AWS demo, not independent testing.

What it means for everyday users and businesses

For ordinary users, this change won't directly alter the apps they use every day. It does reflect AI agents moving from "one question, one answer" toward "multiple agents dividing labor to complete long tasks over several days". AWS says the architecture is not specific to music. It can be applied to GPU-dependent work such as 3D rendering, simulation, model inference and media processing.

For enterprise teams, AWS highlights several benefits. Each team can update its own agent without affecting others. Containers and code packages, two common ways of packaging software, can coexist on the same infrastructure. Sessions can also be stopped when work pauses and resumed later. However, these advantages currently rest on AWS's own claims and demo.

What remains to be seen

  • Information so far comes only from the official AWS blog; performance and cost have not been independently verified by third parties.
  • The AWS post does not give an official launch date for Runtime Instances or a complete list of regional availability.
  • In the demo, catalog comparison was used only against the studio's own catalog; its effectiveness for real-world copyright review remains to be seen.

Frequently asked questions

What are Runtime Instances?

According to AWS, they are a new compute option for hosting AI agents on Amazon Bedrock AgentCore. They are built on AWS-managed EC2 cloud servers and suit persistent, long-running agent workflows. They use the same runtime API as serverless MicroVMs.

What is the biggest difference from the existing MicroVMs?

Per AWS's comparison, MicroVM sessions last up to 8 hours, each microVM hosts one agent, and GPUs are not supported. Runtime Instances sessions last up to 14 days and one instance can host multiple agents. GPUs and EBS persistent storage are also available on supported instance families.

How do multiple agents collaborate on the same machine?

AWS explains that this happens when multiple agent runtimes share the same capacity provider and are invoked with the same runtimeSessionId. They are then placed on the same EC2 instance and share a file system, so they can read files the others produce.

Will I still be charged after stopping?

AWS says that after calling StopRuntimeSession, the instance automatically goes idle and incurs no compute charges while idle. To avoid ongoing charges, AWS recommends deleting the session first. This releases the instance, network interfaces and EBS volumes.

Is this only for music?

No. AWS says music is just a convenient demo subject. The same three-agent architecture can be used for GPU workloads such as 3D rendering, simulation, model inference and media processing.

Are these performance figures reliable?

The figures in the article, for example generating 20 seconds of audio in about 9 seconds, come from a single AWS demo. They are AWS's own claims and have not been independently verified. Actual performance may vary with configuration and environment.

Browse the latest news in this topic

Latest travel guides

Sources

Lifestyle