Lifestyle
AWS introduces Bedrock AgentCore Runtime Instances: AI agents can run for 14 days, use GPUs and collaborate on one machine
In a 2026-09-30 post, AWS used a music production pipeline to demonstrate Runtime Instances, a new way to host AI agents on Amazon Bedrock AgentCore. It offers sessions of up to 14 days, GPU access, lasting storage and several agents working together on one machine. It matters mainly to developers and businesses building AI agents. All information comes from a single source: AWS.
About 7 min read

What happened
On 2026-09-30, AWS published a post on its Artificial Intelligence blog explaining how to build a multi-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances. An AI agent is software that uses an AI model to carry out a task step by step. According to AWS, AgentCore now offers two compute options for hosting AI agents. The first is serverless MicroVMs, where AWS runs small, short-lived machines for you. The second is Runtime Instances, which AWS calls the "new option". Runtime Instances use AWS-managed EC2 infrastructure (EC2 is Amazon's rentable cloud servers) and are aimed at persistent, long-running agent workflows.
AWS notes that as organizations move from single-purpose agents to multi-agent systems, infrastructure needs change. If multiple agents must share context across work spanning several days, serverless sessions capped at a few hours are not enough. The post does not give an official launch date for Runtime Instances, so the event date here is the post's publication date.
Read the full description
Sources are collected, independently checked, then reviewed by Jev.
How MicroVMs and Runtime Instances differ
AWS says both options support agent-building frameworks such as CrewAI, LangGraph, LlamaIndex and Strands Agents. Both also let you choose your foundation model, the underlying AI model, and integrate with MCP and A2A, which are protocols for connecting agents to tools and to each other. The main difference lies in the underlying compute model. The table below summarizes the comparison in the AWS post.
| Item | MicroVM (serverless) | Runtime Instances |
|---|---|---|
| Compute | Fully managed by AWS | AWS-managed EC2 instances |
| Session length | Up to 8 hours | Up to 14 days |
| Agents per compute unit | One agent per microVM (1:1) | Multiple agents per EC2 instance (1:N) |
| Artifact types | Container images and Amazon S3 sources | Container images and Amazon S3 sources |
| GPU | Not supported | Supported (on supported instance families only) |
| Persistence | Session-scoped | Amazon EBS persistent storage |
| Pricing | Pay per use | EC2 runs in your account; Savings Plans and ODCRs can be used |
| Scaling | On demand | Managed by a capacity provider |
According to AWS, the key mechanism is the "shared session". Each agent is deployed as its own runtime. A capacity provider is the setting that tells AgentCore which machines to set up. When two agent runtimes use the same capacity provider and are invoked with the same runtimeSessionId (a session identifier), AgentCore places them on the same EC2 instance. There they share a file system and can read each other's output directly.
AWS's demo: three agents collaborate on a track
- Composer agent: according to AWS, it uses Claude Sonnet 4.6 to turn the producer's request into a music brief. It then generates audio on the instance's GPU with ACE-Step, an open-source music generation foundation model.
- Delivery agent: reads and measures the track on the shared file system. It then asks Claude Sonnet 4.6 to plan EQ, compression and limiting (standard audio adjustments to tone and volume) based on the measurements. It applies them and measures again to confirm the targets are met.
- Compliance agent: independently re-measures the finished track and checks the delivery targets. It also compares harmonic similarity against the studio's own catalog. If it finds a similarity, it calls back the composer agent to produce an alternative version.
AWS published sample results on a g6.xlarge instance in the us-east-2 region. Preparing the model stack took 239 seconds. Composing took 25 seconds, including 8.98 seconds of rendering on an NVIDIA L4 with peak VRAM (GPU memory) of 7.63 GiB. Delivery took 41 seconds and compliance screening took 28 seconds, with all five steps completed on the same instance. Before delivery processing the audio measured -7.5 LUFS (a loudness measure) with a 0.42 dBTP peak (the highest signal level). Afterward it measured -14.0 LUFS with a -3.2 dBTP peak. The compliance screening result was "REVIEW REQUIRED" (human review needed). These figures come from a single AWS demo, not independent testing.
What it means for everyday users and businesses
For ordinary users, this change won't directly alter the apps they use every day. It does reflect AI agents moving from "one question, one answer" toward "multiple agents dividing labor to complete long tasks over several days". AWS says the architecture is not specific to music. It can be applied to GPU-dependent work such as 3D rendering, simulation, model inference and media processing.
For enterprise teams, AWS highlights several benefits. Each team can update its own agent without affecting others. Containers and code packages, two common ways of packaging software, can coexist on the same infrastructure. Sessions can also be stopped when work pauses and resumed later. However, these advantages currently rest on AWS's own claims and demo.
What remains to be seen
- Information so far comes only from the official AWS blog; performance and cost have not been independently verified by third parties.
- The AWS post does not give an official launch date for Runtime Instances or a complete list of regional availability.
- In the demo, catalog comparison was used only against the studio's own catalog; its effectiveness for real-world copyright review remains to be seen.
Frequently asked questions
What are Runtime Instances?
According to AWS, they are a new compute option for hosting AI agents on Amazon Bedrock AgentCore. They are built on AWS-managed EC2 cloud servers and suit persistent, long-running agent workflows. They use the same runtime API as serverless MicroVMs.
What is the biggest difference from the existing MicroVMs?
Per AWS's comparison, MicroVM sessions last up to 8 hours, each microVM hosts one agent, and GPUs are not supported. Runtime Instances sessions last up to 14 days and one instance can host multiple agents. GPUs and EBS persistent storage are also available on supported instance families.
How do multiple agents collaborate on the same machine?
AWS explains that this happens when multiple agent runtimes share the same capacity provider and are invoked with the same runtimeSessionId. They are then placed on the same EC2 instance and share a file system, so they can read files the others produce.
Will I still be charged after stopping?
AWS says that after calling StopRuntimeSession, the instance automatically goes idle and incurs no compute charges while idle. To avoid ongoing charges, AWS recommends deleting the session first. This releases the instance, network interfaces and EBS volumes.
Is this only for music?
No. AWS says music is just a convenient demo subject. The same three-agent architecture can be used for GPU workloads such as 3D rendering, simulation, model inference and media processing.
Are these performance figures reliable?
The figures in the article, for example generating 20 seconds of audio in about 9 seconds, come from a single AWS demo. They are AWS's own claims and have not been independently verified. Actual performance may vary with configuration and environment.
Browse the latest news in this topic
Lifestyle
Google Cloud Launches Spanner Queues: Putting Message Queues Inside Database Transactions to Make AI Agents More Reliable
Google Cloud has announced the general availability of Spanner queues, which make message creation part of a database transaction. The aim is to stop AI agents' "state" and "actions" from falling out of sync. This article covers Google Cloud's claims, the main features, and what it means for general readers.
Lifestyle
GPT-6.1 Sol Launches: New Sol Version in the API, Codex and ChatGPT Work, Not in Chat
OpenAI launched GPT-6.1 Sol on September 29, 2026, with the API name gpt-6.1-sol. The launch rollout covers Codex and ChatGPT Work on Plus, Pro, Business, Enterprise and Edu (Enterprise and Edu need an administrator to enable it); Free and Go are not included at launch, and it is not in Chat (checked September 2026).
Lifestyle
Claude Sonnet 5.5 Launches: Same List Price as Sonnet 5, Available in the API, on Cloud Platforms and in Claude.ai
Anthropic launched Claude Sonnet 5.5 on September 28, 2026. API list prices are the same as Sonnet 5 ($2 per million input tokens, $10 per million output tokens). It is available in Claude.ai, the API and several cloud platforms, and higher-risk cybersecurity requests fall back to Sonnet 5 (checked September 2026).
Lifestyle
Claude Code Adds mods: TypeScript Functions in Plugins Change Its Behavior and Interface
On October 1, 2026, Anthropic introduced Claude Code mods: TypeScript or JavaScript functions that ship inside a plugin and run in the Claude Code process, and can rewrite prompts, manage tool calls and draw new interface. Anthropic says they work in the CLI and the desktop app and are not sandboxed; the documentation says v2.1.287 or later is required (checked October 2026).
Latest travel guides

GuideTokyo
Where to Stay in Tokyo: Comparing Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza, Plus Airport Access, Accommodation Tax, and Luggage Delivery
Where should you stay in Tokyo? Compare Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza by the same criteria: access from Narita and Haneda, transit routes, nearby attractions, neighborhood character, and who each area suits. Includes a comparison table, a Yamanote Line diagram, Tokyo’s accommodation tax as verified in 2026/9 (changing to 3% in 2027/4), and Airport TA-Q-BIN luggage shipping rules.
- Budget
- Hotels

GuideTokyo
How to Choose Tokyo Transit Passes: Are Suica, Welcome Suica, the Tokyo Subway Ticket, and the JR Pass Worth It?
On a first Tokyo trip, start with an IC card and pay per ride (Welcome Suica has no deposit and is valid for 28 days). If you take four or more subway rides in a day, add a 72-hour Tokyo Subway Ticket for 2,000 yen; a JR Pass is never worthwhile if you stay in Tokyo and do not go to Kansai. See what TOURIST PASMO, Suica on iPhone, and the Tokyo Metro day pass do and do not cover, with a decision chart. Prices verified in September 2026.
- Transport
- Budget

GuideTokyo
Tokyo Disneyland and DisneySea Guide: Ticket Prices, Fantasy Springs, Disney Premier Access (DPA), Standby Pass, and Which Park to Choose for Your First Visit
Tokyo Disney one-day Passport prices vary: most weekdays in 9/2026 cost ¥9,900 and weekends ¥10,900. At 14:00 daily, tickets go on sale for the same date two months later. Free Priority Pass is no longer on the official service list; only paid Disney Premier Access (¥1,000–3,500 per person per use) shortens waits. Covers hours, the 25th anniversary, Standby Pass, Entry Request, Fantasy Springs access and first-visit park choice; checked on the official site in 9/2026.
- Itineraries
- Family