Skip to content

MODULE 31 — Cloud Provider AI Platforms: AWS, Microsoft & Google Cloud

Currency note. Accurate as of October 2026. Platform features and names change frequently: all three clouds renamed or restructured their AI platforms within roughly twelve months (Azure AI Foundry became Microsoft Foundry; Vertex AI became the Gemini Enterprise Agent Platform; Bedrock Agents became Bedrock Agents Classic). Treat every product name, limit, count, and price in this module as a dated snapshot, and verify against the provider's current documentation before committing. Volatile facts are tracked in Appendix G §G.6. The decision logic in this module (ecosystem fit first; native infrastructure, portable logic) changes far more slowly than the product names.

31.1 Why This Module Exists

Modules 1-30 are deliberately vendor-neutral. The patterns — RAG, agents, evaluation, governance — apply regardless of which cloud provider you use. But when you deploy those patterns, you deploy them on actual infrastructure from actual providers. This module maps the vendor-neutral patterns to the provider-specific implementations.

The goal is not to pick a winner. Each provider has genuine strengths and genuine weaknesses. The goal is to help architects make rational, evidence-based platform decisions — and to understand the architectural implications of each choice.

The three platforms covered: AWS Bedrock + AgentCore, Microsoft Foundry (formerly Azure AI Foundry), and Google's Gemini Enterprise Agent Platform (formerly Vertex AI). These are the enterprise platforms architects work with. Consumer APIs (direct Anthropic, OpenAI, Google API) are covered in Module 2. This module covers the managed enterprise layer.

What changed in 2026, and why it matters for this module: the model no longer dictates the cloud for the major frontier models. As of October 2026, AWS hosts both Claude and OpenAI models, Microsoft Foundry hosts both OpenAI models and Claude, and Google Cloud hosts Gemini and Claude. Earlier versions of this guidance said "need Claude, pick AWS" or "need GPT, pick Azure." That shortcut no longer holds. Platform choice is now driven by ecosystem fit (identity, data gravity, compliance boundary, team skills), with model availability as a check, not the deciding factor.


31.2 Amazon Bedrock and AgentCore

The Bedrock Architecture

Amazon Bedrock is AWS's managed foundation model service — it provides access to foundation models from Anthropic (Claude), OpenAI, Meta (Llama), Mistral, Amazon (Titan/Nova), and others, all within your AWS account. Data stays in your AWS region.

Bedrock's model catalog is curated rather than comprehensive — roughly 30 production-tested, regionally available models in early 2025, grown to nearly 100 serverless models plus 100+ more via Bedrock Marketplace by mid-2026 (Anthropic, OpenAI, Meta, Mistral, Cohere, AI21, Amazon Titan/Nova, DeepSeek, Google Gemma, OpenAI's open-weight gpt-oss models, Qwen3, Stability AI, and others). The catalog is smaller than Microsoft Foundry's, which advertises over ten thousand models, but every serverless model on Bedrock has passed AWS's production validation — fewer but more vetted. Check the current count before quoting it; this list grows every quarter.

OpenAI models on Bedrock (2026). OpenAI's frontier models became generally available on Bedrock in stages: GPT-5.5, GPT-5.4, and Codex in June 2026; the GPT-5.6 family (Sol, Terra, Luna) in July 2026; and GPT-6 Astra, GPT-6 Sol/Luna, and GPT-6.1 Sol in September 2026. For the June launch, AWS stated that GPT-5.5 and GPT-5.4 pricing matched OpenAI's first-party rates and that usage counts toward existing AWS commitments. Do not generalize that to every model or later release: check pricing per model (see "Same Model, Different Clouds" in §31.5). OpenAI's open-weight gpt-oss and NVIDIA Nemotron models are also available in AWS GovCloud (April 2026).

BEDROCK CORE SERVICES

Amazon Bedrock (foundation layer):
  ├── Foundation model access (Claude, OpenAI, Llama, Mistral, Titan, etc.)
  ├── Bedrock Guardrails — content filtering for input/output
  │     Blocks up to 88% of harmful outputs (AWS-published benchmark)
  │     Includes automated reasoning checks for factual accuracy
  ├── Bedrock Knowledge Bases — managed RAG
  │     Handles: ingestion, chunking, embedding, vector storage (OpenSearch)
  │     You supply: documents. Bedrock does the pipeline.
  ├── Bedrock Intelligent Prompt Routing — automatic model routing
  │     Routes between models (e.g., Claude Sonnet ↔ Claude Haiku)
  │     based on query complexity for cost/performance optimization
  └── Bedrock Evaluations — model and prompt evaluation service

Amazon Bedrock Agents Classic (the original agent framework, Nov 2023):
  ├── Action groups — Lambda functions as agent tools
  ├── Knowledge base integration — RAG connected to agents
  ├── Multi-agent collaboration — orchestrator + subagent patterns
  └── Status: maintenance mode. Closed to accounts without prior usage
        from July 30, 2026; no new features and a frozen model catalog;
        existing agents keep running; no announced end-of-life date.
        AWS recommends AgentCore for new work.

Amazon Bedrock AgentCore (GA: October 2025)

AgentCore is the enterprise-grade successor to Bedrock Agents — a fully managed, modular platform for deploying production AI agents at scale. The key differentiator from Bedrock Agents: AgentCore is framework-agnostic and model-agnostic. It works with LangGraph, CrewAI, LlamaIndex, Strands Agents, and other frameworks, and with models inside or outside Bedrock.

With AgentCore, you can enable agents to take actions across tools and data with the right permissions and governance, run agents securely at scale, and monitor agent performance and quality in production — all without any infrastructure management.

AGENTCORE COMPONENT MAP (as of October 2026)

AgentCore Runtime:
  What it does: Managed execution environment for AI agents
  Two execution options:
    ├── Default serverless microVM sessions: up to 8 hours, fast startup,
    │     session isolation
    └── Runtime instances (dedicated EC2-backed compute in your account;
          GA August 2026): sessions up to 14 days; GPU-accelerated,
          memory-optimized, and compute-optimized families
  Networking: VPC support, PrivateLink
  A2A support: Agent-to-Agent protocol supported (since October 2025)
  Framework compatibility: LangGraph, CrewAI, LlamaIndex, Strands Agents,
                           OpenAI Agents SDK, Claude Agent SDK, custom
  Architecture pattern: Deploy your agent code as a container;
                        AgentCore handles scaling, session isolation,
                        and operational concerns

AgentCore Managed Harness (GA June 2026):
  What it does: Config-based agent — declare model, tools, and
                instructions; AgentCore runs the loop, compute, memory,
                identity, and observability (no container by default)
  Closest analog to Bedrock Agents Classic; the recommended migration
  target. For custom orchestration or multi-agent routing, deploy
  code-defined agents on Runtime instead.
  Harness choice in general: see Module 6 §6.5.

AgentCore Gateway:
  What it does: Converts APIs, Lambda functions, and MCP servers into
                agent-accessible tools; single discovery endpoint
  Key capability: Any existing Lambda function becomes an MCP-compatible
                  tool with minimal code changes
  Auth: IAM authorization + OAuth for secure agent-to-tool interactions
  Architecture pattern: The MCP server management problem (Module 7)
                        solved as a managed service

AgentCore Memory:
  What it does: Managed short-term and long-term memory for agents
  Memory types: Session memory (within conversation) + persistent memory
                (across sessions, with episodic learning capability)
  Self-managed option: Control your own memory extraction/consolidation
  Architecture pattern: The persistent memory patterns from Module 21
                        without the infrastructure management

AgentCore Identity:
  What it does: Managed identity for agents — workload identities with
                outbound auth to third-party services and AWS services
  Approach: IAM-integrated, plus OAuth2 and API-key credential providers
  Architecture pattern: Agent authorization without shared credentials

AgentCore Policy (GA March 2026):
  What it does: Policy enforcement over agent tool access at the Gateway

AgentCore Evaluations (GA March 2026):
  What it does: Managed quality evaluation of agent behavior

AWS Agent Registry (own service namespace since August 2026):
  What it does: Catalog of approved agents, MCP servers, and skills with
                an approval workflow
  Migration note: registry usage under the old bedrock-agentcore namespace
                  must be migrated; the old namespace is scheduled to
                  shut down October 30, 2026

AgentCore Observability:
  What it does: Real-time visibility into agent execution and metrics
  Stack: Amazon CloudWatch + OpenTelemetry compatible
  Covers: End-to-end trace, tool call logs, memory operations,
          latency, cost per agent task

AgentCore Code Interpreter:
  What it does: Sandboxed code execution environment for agents
  Languages: Multiple (Python, JavaScript, others)
  Security: Isolated sandbox per execution

AgentCore Browser Tool:
  What it does: Cloud-hosted browser runtime for web-browsing agents
  Use case: Computer use patterns (Module 21) as managed infrastructure

When to Choose AWS / Bedrock / AgentCore

AWS IS THE RIGHT CHOICE WHEN:
  ├── Your organization is already deeply embedded in AWS
  │     (IAM, VPC, Lambda, S3, CloudWatch are your operational baseline)
  ├── You want Claude and OpenAI frontier models under one AWS contract,
  │     IAM boundary, and bill (verify per-model availability and pricing)
  ├── Your compliance story relies on AWS certifications
  │     (FedRAMP, SOC2, HIPAA on AWS)
  ├── You want curated model quality over model breadth
  └── You need very long agent sessions on infrastructure you control
        (runtime instances: up to 14 days, in your account)

AWS IS NOT THE BEST CHOICE WHEN:
  ├── Your organization runs primarily on Microsoft 365 and Azure
  ├── You need the broadest model catalog (Foundry advertises 10,000+)
  ├── Your agentic workflows need to integrate deeply with M365
  └── You need a model family that Bedrock does not host
        (check the current catalog; e.g., Gemini is Google-hosted)

KEY AWS-SPECIFIC ARCHITECTURAL DECISION:
  Bedrock Agents Classic vs. AgentCore:
  ├── Bedrock Agents Classic: in maintenance mode; new accounts cannot
  │     create agents since July 30, 2026
  ├── AgentCore harness: config-based, closest to Classic; start here
  │     for simple agents
  ├── AgentCore Runtime with code-defined agents: framework-agnostic,
  │     production-grade ops; for complex multi-agent, custom
  │     orchestration, or long-running workflows
  └── Migration path: Classic → AgentCore harness, then to code-defined
        agents as complexity grows

31.3 Microsoft Foundry (formerly Azure AI Foundry)

The Foundry Architecture

Microsoft Foundry is the current name for the platform previously called Azure AI Foundry (and before that Azure AI Studio). The rename was announced at Ignite in November 2025 and became official in Microsoft's Product Terms in January 2026. It folds Azure AI Foundry, Azure AI Studio, and Azure AI Services into one resource and portal. It remains the dominant choice for Microsoft-first enterprises — the platform with the tightest integration into M365, Azure, and the Microsoft security ecosystem.

Microsoft Foundry is positioned as a unified enterprise platform for model access, agent orchestration, and governed deployment. It is designed to reduce architectural fragmentation and help cross-functional teams collaborate in one environment.

Structure change to know about. The new Foundry portal is hub-less: projects sit directly under a Foundry resource. The older hub-based project model (an AI Hub plus separately managed storage, container registry, and Key Vault) lives on in the "Foundry (classic)" portal. Microsoft's documentation states that new agent and model-centric capabilities land on the new Foundry project type, including Foundry Agent Service. If your team learned the hub-and-project model, plan a migration review rather than assuming it carries forward.

MICROSOFT FOUNDRY COMPONENT MAP (as of October 2026)

Foundry Resource and Projects:
  What it does: One resource with projects beneath it (new portal);
                projects are the container for access management,
                data, and monitoring
  Classic: hub-based projects remain in the classic portal
  Architecture value: Centralized governance with team-level isolation

Foundry Model Catalog:
  Breadth: 11,000+ models by Microsoft's marketing figure (technical
           documentation says "over 10,000"); includes OpenAI models,
           Claude, Llama, Mistral, DeepSeek, Cohere, Phi (Microsoft
           first-party), and many others
  OpenAI models: served under your Azure tenant and enterprise terms
  Claude: available in the catalog; billed at Anthropic first-party rates
  Advantage over Bedrock: Much broader; access to more open-weight models
  Tradeoff: Less curation; production-readiness varies by model

Foundry Agent Service (GA March 2026):
  What it does: Managed agent execution within Foundry
  Hosted Agents (GA July 9, 2026): bring agent code built with your
    preferred framework; runs in Foundry's managed runtime. Each session
    is a stateful, isolated sandbox with a persisted filesystem.
    ├── Sessions persist for up to 30 days
    ├── Idle timeout 2-60 minutes (default 15); compute is deprovisioned
    │     and state restored when the session is referenced again
    └── A maximum continuous execution duration is not stated in the
          documentation reviewed for this edition. Do not equate the
          30-day session lifetime with 30 days of continuous execution.
  Tools: MCP servers and A2A endpoints can be connected as tools, with
         OAuth passthrough, Entra Agent Identity, or managed identity
         for authentication
  Integration: Microsoft 365 surfaces (Teams, Outlook, SharePoint)

Microsoft Agent Framework (1.0 released April 3, 2026):
  What it does: Open-source SDK unifying Semantic Kernel and AutoGen;
                supports MCP and A2A; Python and .NET
  Role: Microsoft's recommended code-first replacement for Prompt Flow
  Harness choice in general: see Module 6 §6.5

Microsoft Entra Agent ID (GA May 2026):
  What it does: Agents are first-class directory objects with their own
                identities (blueprint → agent identity → optional agent user)
  Implications:
    ├── Agent sign-in activity appears in Entra audit logs
    ├── Conditional Access policies apply to agents (including templates
    │     for autonomous and on-behalf-of agents)
    ├── Agent access can be governed with access packages and sponsors
    │     (lifecycle management like human identities)
    └── Microsoft documents integration for non-Microsoft agents,
          including a guide for securing Amazon Bedrock agents
  Architecture significance: For organizations already managing human
    identity through Entra ID, this is the cleanest governance integration.
    Agent registry experiences are converging under Microsoft Agent 365.

Prompt Flow (RETIRING):
  Status: feature development ended April 20, 2026; retirement April 20,
          2027. Not recommended for new development.
  Migration: Microsoft Agent Framework (code-first; no visual editor)
  Implication: if your team relied on the visual designer, plan for a
               move to code-defined workflows

Foundry IQ and Azure AI Search (RAG layer):
  What it does: Azure AI Search provides enterprise search with vector +
                semantic + hybrid search. Foundry IQ knowledge bases (GA)
                build on it to ground agents in enterprise data, with an
                MCP server for any MCP-compatible host.
  Integration: SharePoint, Microsoft 365, custom data sources

When to Choose Azure / Foundry

AZURE IS THE RIGHT CHOICE WHEN:
  ├── Your organization runs on Microsoft 365
  │     (Teams, SharePoint, Exchange, Outlook integration is built-in)
  ├── Entra ID is your identity foundation
  │     (Entra Agent ID gives the cleanest agent governance story)
  ├── Your compliance boundary is Azure
  │     (Broadest compliance certifications across regulated industries)
  ├── You need the broadest model catalog, including Microsoft first-party
  │     models, alongside OpenAI and Claude
  └── Copilot Studio / M365 Copilot is already in your environment

AZURE IS NOT THE BEST CHOICE WHEN:
  ├── Your infrastructure is AWS-primary
  ├── You need a published, very long continuous execution window on
  │     infrastructure you control (AWS runtime instances: 14 days;
  │     Google Agent Runtime: 7 days; verify Foundry's current limits)
  ├── Your data pipeline is GCP-native (BigQuery, Cloud Storage)
  └── You want to avoid Microsoft ecosystem dependency

KEY AZURE-SPECIFIC ARCHITECTURAL DECISION:
  Microsoft Agent Framework vs. LangGraph (or other framework):
  ├── Agent Framework: Microsoft-backed, open source, MCP + A2A, first-class
  │     Foundry hosting and tracing, the supported path off Prompt Flow
  └── LangGraph/custom: more portable across clouds, mature checkpointing;
       you own more of the integration

  Either way, keep business logic in portable code. Prompt Flow's
  retirement is a live example of why (see §31.10, Anti-Pattern 1).

  Entra Agent ID is the most developed per-agent directory identity of the
  three clouds. AWS (AgentCore Identity workload identities) and Google
  (Agent Identity) now have native agent identities too, so the difference
  is integration with your existing human-identity system, not existence.

31.4 Google Gemini Enterprise Agent Platform (formerly Vertex AI)

The Platform Architecture

On April 22, 2026, Google announced the Gemini Enterprise Agent Platform as the replacement for Vertex AI; Google stated that Vertex AI services and roadmap evolution are now delivered through Agent Platform. It is the platform of choice for data-heavy organizations already running on GCP — where BigQuery, Cloud Storage, and Google's data infrastructure are the foundation.

Naming caution. Several sub-products were renamed with the platform. Names below marked "(reported)" come from secondary sources and should be checked against Google's release notes. Names that I could not verify keep their Vertex-era form.

GEMINI ENTERPRISE AGENT PLATFORM COMPONENT MAP (as of October 2026)

Model Garden:
  Models: Gemini family (primary), Claude, open-weight models, and
          partner models (a figure of 200+ is reported)
  Gemini advantage: Native Gemini access with best model version support
  Current Gemini models: Gemini 3.8 Flash (GA; 1M-token input context),
          Gemini 3.1 Pro (preview as of October 2026)
  Former name: Vertex AI Model Garden (reported)

Agent Studio:
  What it does: Low-code environment for prompt design and agent building
  Former name: Vertex AI Studio (reported)

Agent Development Kit (ADK) and Agent Builder (Vertex-era name; verify
current naming):
  What it does: Code-first and low-code routes to build agents and RAG
                systems; Dialogflow CX integration for voice/telephony
                agents (Vertex-era integration; verify current status)
  Playbooks: natural-language agent instructions (Vertex-era feature;
             verify before relying on it)
  Grounding: Google Search grounding for web-current information

Agent Runtime (reported: formerly Agent Engine):
  What it does: Managed runtime for production agents
  Key spec: agents that run continuously for up to 7 days (Google blog)
  Related services: Memory Bank (reported: renamed to Agent Platform
                    Memory Bank), Sessions

Governance layer (Google blog):
  ├── Agent Identity — native IAM type for agents, built on open
  │     standards, least-privilege permissions
  ├── Agent Gateway — central control point securing and governing
  │     user-agent, agent-tool, and agent-agent interactions
  ├── Agent Registry — single library of agents, servers, connections
  ├── Agent Evaluation — online evaluation monitors in production
  └── Agent Observability — traces of reasoning and tool use

Vertex AI Search (Vertex-era name; verify current naming):
  What it does: Enterprise search with semantic understanding
  Grounding options: Google Search, enterprise data, custom corpora
  Unique capability: Google Search grounding (real-time web information
                     in enterprise RAG responses)

Vertex AI Pipelines (Vertex-era name; verify current naming):
  What it does: ML workflow orchestration for training and evaluation
  Best for: Organizations running ML training alongside inference

MCP and A2A:
  Google originated A2A and supports it natively. Google also published
  a remote MCP server for the Agent Platform, so external agents (for
  example coding agents) can reach platform resources over MCP.

Gemini for Google Workspace:
  What it does: AI assistance across Gmail, Docs, Sheets, Meet, Drive
  Equivalent to: Microsoft 365 Copilot (on the Google side)

When to Choose GCP / Gemini Enterprise Agent Platform

GCP IS THE RIGHT CHOICE WHEN:
  ├── Your data infrastructure is already on GCP
  │     (BigQuery, Cloud Storage, Pub/Sub are your data foundation)
  ├── Google Search grounding is architecturally important
  │     (Web-current RAG responses require Google's search index)
  ├── Gemini is your preferred model family (Claude is also available)
  ├── You want native A2A support from the protocol's originator
  └── Google Workspace (not M365) is your productivity stack

GCP IS NOT THE BEST CHOICE WHEN:
  ├── Your infrastructure is AWS or Azure primary
  ├── You need the deepest enterprise compliance certifications
  ├── Your team doesn't have GCP expertise
  ├── You need OpenAI frontier models in the same boundary (not verified
  │     on Google Cloud as of this edition)
  └── Your agent workflows need M365 or AWS-native integrations

31.5 The Platform Comparison Matrix

ENTERPRISE AI PLATFORM COMPARISON (as of October 2026)

                       AWS Bedrock/AgentCore  Microsoft Foundry   Gemini Ent. Agent
                                                                  Platform (ex-Vertex)
──────────────────────────────────────────────────────────────────────────────────────
MODEL ACCESS (the model no longer dictates the cloud for frontier models)
  Catalog breadth      Curated (~100          11,000+ (company    Model Garden
                       serverless + 100+      figure; docs say    (200+ reported)
                       Marketplace, mid-2026) 10,000+)
  OpenAI frontier      Yes (GA Jun-Sep 2026)  Yes                 Not verified
  Claude               Yes                    Yes (Anthropic      Yes
                                              first-party rates)
  Gemini               No (Gemma open-weight) Not verified        Yes (primary)
  Open-weight models   Yes (curated)          Extensive           Yes
  Pricing vs. direct   Verify per model       Claude: first-party Verify per model
  API                  (see below)            rates               (regional premium
                                                                  on some endpoints)

AGENT INFRASTRUCTURE
  Managed runtime      AgentCore Runtime +    Agent Service +     Agent Runtime
                       managed harness        Hosted Agents
  Longest published    8 h (microVM);         Sessions persist    Up to 7 days
  execution window     14 days (runtime       up to 30 days;
  (Oct 2026)           instances, GA Aug      max continuous
                       2026)                  run not published
  Framework agnostic   Yes (LangGraph,        Agent Framework     ADK + others
                       Strands, OpenAI SDK,   (open source) +     (verify)
                       Claude Agent SDK)      bring-your-own code
  A2A support          Yes                    A2A tools; Agent    Yes (originated)
                                              Framework (MCP+A2A)
  MCP support          AgentCore Gateway      MCP tools in Agent  Agent Gateway;
                                              Service; Foundry IQ remote MCP server
                                              MCP server
  Agent registry       AWS Agent Registry     Agent 365 registry  Agent Registry
                                              (converging)

IDENTITY & SECURITY
  Agent identity       AgentCore Identity     Entra Agent ID      Agent Identity
                       (workload identities;  (GA May 2026;       (native IAM type)
                       IAM + OAuth)           directory objects)
  Agent policy         AgentCore Policy       Conditional Access  Agent Gateway
                       (GA Mar 2026)          for agents          governance
  Audit trail          CloudTrail/CloudWatch  Entra ID audit log  Cloud Audit Logs
  Private networking   VPC + PrivateLink      VNet + Private Link VPC + Private SC

MANAGED RAG
  Service name         Bedrock Knowledge Base Foundry IQ + Azure  Vertex AI Search
                                              AI Search           (Vertex-era name)
  Web grounding        Not native             Bing (verify)       Google Search ✓
  Hybrid search        Yes                    Yes                 Yes

EVALUATION & OBSERVABILITY
  Agent evaluation     AgentCore Evaluations  Foundry evaluations Agent Evaluation
                       (GA Mar 2026)
  Native stack         CloudWatch + OTel      Azure Monitor + OTel Cloud Ops + OTel
  OTel compatible      Yes ✓                  Yes ✓               Yes ✓

ECOSYSTEM FIT
  Best for             AWS-primary orgs       Microsoft 365 orgs  GCP-primary orgs
  M365 integration     Poor                   Excellent ✓         Poor
  Google Workspace     Poor                   Poor                Excellent ✓
  AWS services         Excellent ✓            Limited             Limited

COMPLIANCE
  FedRAMP              Yes                    Yes                 Yes
  Regulated industry   Strong (AWS certs)     Strongest breadth ✓ Good
  EU data residency    Available              Available           Available

ARCHITECTURAL CEILING
  Typical strength     Claude + OpenAI under  M365 integration,   Data-heavy GCP
                       one IAM boundary,      Entra agent         workloads,
                       very long sessions,    governance, widest  Google Search,
                       AWS ecosystem          catalog             Gemini, A2A

  Typical complaint    "Ship slower than      "Fragmented         "Less enterprise
                       outside Bedrock"       observability;      mature than AWS";
                                              frequent renames"   frequent renames

Where the matrix says "Not verified", the claim was not confirmed against provider documentation for this edition. It is not evidence the capability is absent.

Same Model, Different Clouds

Because the major frontier model families now appear on more than one cloud, "which model" and "which cloud" are separate decisions. That creates two options that did not exist a year ago:

  • Portability: run the same model family on whichever cloud already holds your data, identity, and compliance boundary, instead of moving your estate to follow the model.
  • Resilience: use the same model on a second cloud as a fallback target, rather than (or in addition to) a different model. The prompts, tool schemas, and evaluation baselines you already validated mostly carry over, so the fallback is cheaper to qualify than a cross-vendor swap.

Same model does not mean same service. Treat each cloud's offering as a distinct deployment target:

SAME MODEL ON A SECOND CLOUD — WHAT TO VERIFY

  ├── Pricing: partner-priced offerings do not necessarily match the
  │     first-party API. At its June 2026 launch AWS stated that GPT-5.5
  │     and GPT-5.4 matched OpenAI's first-party rates; Anthropic
  │     documents a 10% premium for regional (versus global) endpoints
  │     of recent Claude models on Bedrock and Vertex-era endpoints;
  │     Claude on Foundry is billed at Anthropic first-party rates.
  │     Verify per model and per endpoint type; do not assume parity.
  ├── Feature lag: new model versions, beta features, tools, and
  │     parameters can arrive on a partner cloud later than on the
  │     first-party API, or in a different form. Test the exact features
  │     your capability uses (tool use, caching, batch, structured output).
  ├── Regional availability: a model offered in one region of one cloud
  │     may be absent in your required region on another. Check the
  │     regions that satisfy your data-residency rules.
  ├── Quotas and throughput: quotas, rate limits, and provisioned-
  │     capacity options are separate per cloud and per account. A
  │     fallback account sized for 5% keep-alive traffic will throttle
  │     if it suddenly takes 100%.
  ├── Data terms: retention, zero-data-retention, and BAA coverage are
  │     per service and per model. Confirm them for the second cloud.
  └── Model identifiers and APIs differ: wrap them behind your adapter
        layer rather than hard-coding either cloud's SDK.

The failover mechanics (hot, warm, and cold readiness, keep-alive traffic, triggers, and the cache-cold cost spike) are covered in Module 37 §37.7; this module does not repeat them. A same-model second-cloud fallback usually sits at the "L1 hot fallback, same quality tier" rung of that module's degradation ladder, and it does not protect against a failure of the model vendor itself. Pair it with a different-model fallback for capabilities where that risk matters. For where the agent loop itself runs and how portable it is across these clouds, see Module 6 §6.5.


31.6 Provider-Native vs. Vendor-Neutral: The Decision

The fundamental architectural question for every organization: should you use provider-native AI services (Bedrock AgentCore, Foundry Agent Service, Gemini Enterprise Agent Platform), or should you build on vendor-neutral foundations (LangGraph + vLLM + LiteLLM + Weaviate) that work across providers?

PROVIDER-NATIVE APPROACH

Advantages:
  ├── Faster time to production — managed infrastructure
  ├── Tighter integration with provider ecosystem
  ├── Less operational overhead — the provider manages the plumbing
  ├── Better compliance posture within that provider's trust boundary
  └── Fewer moving parts to govern and monitor

Disadvantages:
  ├── Provider lock-in — migrating is expensive
  ├── Feature releases on provider's schedule, not yours
  ├── Product names, APIs, and namespaces change (three major renames
  │     in twelve months across the big three)
  ├── Less flexibility for complex custom patterns
  └── Cost premium vs. self-managed open source

VENDOR-NEUTRAL APPROACH

Advantages:
  ├── Portability across providers and models
  ├── More control over every component
  ├── Can optimize cost at each layer independently
  └── No lock-in — each component can be replaced

Disadvantages:
  ├── More engineering investment to build and operate
  ├── Integration complexity (each component is a separate system)
  ├── Operational overhead (your team manages what the provider would)
  └── Compliance requires assembling the controls yourself

THE REALISTIC ANSWER FOR MOST ENTERPRISES:
  Start provider-native, selectively go vendor-neutral where it matters.

  ├── Provider-native for: managed RAG, agent hosting infrastructure,
  │     observability (the undifferentiated heavy lifting)
  └── Vendor-neutral for: the logic layer (your LangGraph agent code
        that contains your business logic should be portable),
        the model interface (LiteLLM gateway abstracts provider APIs),
        evaluation (RAGAS/PromptFoo are provider-agnostic)

  The goal: your business logic is portable even if your infrastructure is not.
  Write agent logic that can be deployed to AgentCore, Foundry, or the
  Gemini Enterprise Agent Platform with only configuration changes —
  not code rewrites. Model-level portability is covered in Module 37.

31.7 Mapping the Course Patterns to Provider Implementations

VENDOR-NEUTRAL PATTERN → PROVIDER IMPLEMENTATIONS

Module 4: RAG Architecture
  ┌─────────────────────┬──────────────────────┬─────────────────────┐
  │ Generic Pattern     │ AWS Implementation   │ Microsoft Impl.     │
  ├─────────────────────┼──────────────────────┼─────────────────────┤
  │ Document ingestion  │ Bedrock Knowledge    │ Foundry IQ +        │
  │ pipeline            │ Base data source     │ Azure AI Search     │
  │ Vector store        │ OpenSearch Serverless│ Azure AI Search     │
  │ Hybrid search       │ KB hybrid search     │ AI Search hybrid    │
  │ Confidence gate     │ Custom Lambda layer  │ Agent Framework     │
  │                     │                      │ conditional edge    │
  └─────────────────────┴──────────────────────┴─────────────────────┘

Module 6: Agentic Architecture
  ┌─────────────────────┬──────────────────────┬─────────────────────┐
  │ Agent loop          │ AgentCore Runtime /  │ Foundry Agent       │
  │                     │ managed harness      │ Service (Hosted)    │
  │ Tool manifest       │ AgentCore Gateway    │ Foundry tool config │
  │ Human approval gate │ Lambda + SNS         │ Teams approval flow │
  │ Agent identity      │ AgentCore Identity   │ Entra Agent ID      │
  │ Agent memory        │ AgentCore Memory     │ Foundry memory      │
  │ Observability       │ AgentCore Obs.       │ Azure Monitor       │
  └─────────────────────┴──────────────────────┴─────────────────────┘

Module 7: Multi-Agent / MCP / A2A
  ┌─────────────────────┬──────────────────────┬─────────────────────┐
  │ MCP server hosting  │ AgentCore Gateway    │ Foundry MCP tools   │
  │ A2A protocol        │ AgentCore Runtime    │ MSFT Agent Framework│
  │ Agent registry      │ AWS Agent Registry   │ Agent 365 registry  │
  └─────────────────────┴──────────────────────┴─────────────────────┘

Module 9: Security
  ┌─────────────────────┬──────────────────────┬─────────────────────┐
  │ Content filtering   │ Bedrock Guardrails   │ Azure Content Safety│
  │ Tool authorization  │ IAM + AgentCore      │ Entra ID + Foundry  │
  │                     │ Policy               │                     │
  │ PII detection       │ Presidio + custom    │ Presidio + custom   │
  │ Audit trail         │ CloudTrail +         │ Entra ID audit +    │
  │                     │ AgentCore Obs.       │ Azure Monitor       │
  └─────────────────────┴──────────────────────┴─────────────────────┘

Module 13: Cost Engineering
  ┌─────────────────────┬──────────────────────┬─────────────────────┐
  │ Model routing       │ Bedrock Intelligent  │ Foundry routing     │
  │                     │ Prompt Routing       │ (manual config)     │
  │ Prompt caching      │ Bedrock cache        │ Azure prompt cache  │
  │ Cost attribution    │ Cost allocation tags │ Azure cost mgmt     │
  └─────────────────────┴──────────────────────┴─────────────────────┘

For the Google column, map the same patterns to Agent Runtime, Agent
Gateway, Agent Identity, Agent Registry, Memory Bank, and Cloud
Observability (confirm current service names in Google's release notes).

31.8 The Bedrock AgentCore Architectural Deep Dive

Since AgentCore is the most architecturally significant 2025 development in managed agent infrastructure and is still poorly covered in older resources, this section goes deeper. The 2026 equivalents on Foundry and the Gemini Enterprise Agent Platform follow the same shape (managed runtime, identity, gateway, memory, evaluation), so most of the lessons transfer.

The Production Problem AgentCore Solves

AgentCore shifts AI agent development by providing enterprise-grade services for runtime, identity, memory, and observability, enabling teams to focus on business logic rather than infrastructure. This offers three key benefits: faster deployment, flexible model and framework choice, and built-in security.

The specific problems it eliminates: - Manual container orchestration for agent sessions - Custom session isolation implementation - Rolling your own memory store with encryption and access control - Building a tool discovery and routing layer from scratch - Assembling agent observability from individual AWS services

The AgentCore Gateway as MCP Hub

AgentCore Gateway connects to existing Model Context Protocol (MCP) servers in addition to transforming APIs and Lambda functions into agent-compatible tools. It supports Identity and Access Management (IAM) authorization, enabling customers to leverage IAM in addition to OAuth for secure agent-to-tool interactions over MCP, and acts as a single, secure endpoint for agents to discover and use tools without the need for custom integrations.

This is architecturally significant: AgentCore Gateway is essentially a managed MCP server registry and proxy. The MCP governance registry problem (Module 7) is solved as a managed service — teams register their Lambda functions and APIs once, and any agent with IAM permissions can discover and use them. Google's Agent Gateway and Registry, and Microsoft's Foundry tool catalog with Agent 365 registry convergence, address the same problem.

Long-Running Execution Windows

Long agent execution is no longer an AWS-only capability. As of October 2026, published figures are:

LONG-RUNNING AGENT EXECUTION (as of October 2026)

AWS AgentCore Runtime:
  microVM sessions ............ up to 8 hours (fast startup, serverless)
  runtime instances ........... up to 14 days (GA August 2026; dedicated
                                EC2-backed compute in your account,
                                including GPU families)
Google Agent Runtime .......... agents up to 7 days (Google blog)
Microsoft Foundry Hosted Agents: sessions persist up to 30 days, with an
                                idle timeout (2-60 min, default 15) that
                                deprovisions compute and restores state.
                                A maximum continuous execution duration is
                                not stated in the documentation reviewed.
AWS Lambda (for contrast) ..... 15 minutes maximum

Compare like with like: "session lifetime", "continuous execution", and "agent run" are different quantities, and each provider words its limit differently. Read the current documentation for the exact definition before sizing a design around a number.

The architectural point is unchanged by any single limit: a long execution window removes the need to decompose a task into sub-15-minute pieces, but it does not remove the need for durable task state. A 14-day session is not a substitute for checkpoints (Module 6 §6.5): runtimes get redeployed, sessions get stopped, and models get swapped. Keep canonical task state in your own store.

This is most relevant for: document analysis workflows (processing a large document corpus), research automation, multi-step data processing with external API calls, and multi-day asynchronous work (Module 6, Pattern 5).


31.9 The Platform Selection Checklist

Use this before committing to a provider's managed AI platform.

PROVIDER PLATFORM SELECTION CHECKLIST

ECOSYSTEM FIT (most important criterion)
  □ What is our primary cloud provider?
  □ Does our identity management use Entra ID (→ Microsoft) or IAM (→ AWS)?
  □ Do our workflows primarily run through M365 (→ Microsoft) or AWS (→ AWS)?
  □ Is our data infrastructure on GCP (→ Gemini Enterprise Agent Platform)?

  Rule: Match the AI platform to your existing ecosystem. Switching
  cloud for AI creates operational fragmentation that costs more
  than any AI platform feature differential.

MODEL REQUIREMENTS (a check, not the deciding factor)
  □ Which model families do we need? Confirm each is offered on the
    candidate cloud, in our required regions, with the features we use.
    (As of October 2026: Claude on all three; OpenAI frontier models on
    AWS and Microsoft Foundry; Gemini on Google Cloud.)
  □ Do we need a model that exists on only one cloud? Then record that
    as a platform dependency and plan a fallback model (Module 37).
  □ Do we need the broadest model selection, including long-tail
    open-weight models? → Microsoft Foundry (11,000+ by Microsoft's figure)
  □ Have we checked pricing for this model on this cloud, rather than
    assuming it matches the first-party API?
  □ Could a second cloud serve as a same-model fallback (§31.5)?

AGENT REQUIREMENTS
  □ How long must a single agent run last? Compare current published
    limits: AWS runtime instances up to 14 days, Google Agent Runtime up
    to 7 days, Foundry Hosted Agents (verify the current limit). Then
    design checkpoints regardless.
  □ Do we need agents deployed into Teams/Outlook? → Microsoft Foundry
  □ Do we need Google Search grounding? → Gemini Enterprise Agent Platform
  □ Do we need IAM-integrated tool discovery? → AgentCore Gateway
  □ Do we need per-agent identity in our existing directory? → Entra Agent ID
    (AgentCore Identity and Google Agent Identity are the native
    equivalents on their clouds)
  □ Which harness do we run on the platform? → Module 6 §6.5

COMPLIANCE REQUIREMENTS
  □ What compliance certifications are required?
    (All three platforms offer FedRAMP, SOC2, HIPAA)
  □ What data residency region?
    (All three platforms offer EU residency)
  □ Are AI Act high-risk documentation requirements in scope?
    (All three have commitments; verify current documentation)
  □ Are retention, ZDR, and BAA terms confirmed for the exact model
    and feature on this cloud?

TEAM CAPABILITY
  □ What cloud does our team know well?
  □ Can we operate the platform without significant retraining?
  □ Are we willing to learn a new orchestration paradigm
    (Agent Framework, AgentCore harness, ADK)?
  □ Who tracks product renames and retirements (Prompt Flow retires
    April 2027; Bedrock Agents Classic is frozen)?

  Rule: A platform the team knows moderately well will outperform
  a technically superior platform the team doesn't know.

31.10 Provider-Native Anti-Patterns

Anti-Pattern 1: Business logic in provider-specific constructs

Writing your agent logic in Bedrock Agents Classic action group syntax, or Foundry Prompt Flow YAML, or Vertex-era Playbook format — instead of in portable code (LangGraph, Python functions). When you need to migrate, the logic migration is expensive. This is no longer hypothetical: Bedrock Agents Classic is in maintenance mode (July 2026) and Prompt Flow retires in April 2027, so teams that encoded logic in those constructs are now migrating on the vendor's schedule.

Fix: Agent logic lives in portable code. Provider services handle the infrastructure. The provider's orchestration service runs your code; your code doesn't depend on the provider's orchestration syntax.

Anti-Pattern 2: Assuming managed = production-ready without governance

Bedrock AgentCore handles infrastructure management. It does not handle your prompt governance, your eval suite, your human oversight design, or your model risk documentation. Teams that use "managed service" as a shortcut for "production ready" are missing most of the governance work.

Fix: The provider-specific checklist supplements, not replaces, the universal pre-deployment checklist (Appendix C).

Anti-Pattern 3: Provider selection before use case definition

"We're an AWS shop so we'll use AgentCore" before defining what the agent needs to do, what model it needs, what compliance requirements apply, and whether a simpler non-provider-native approach would work.

Fix: Platform selection is the output of use case analysis, not the input. Define requirements first. Select platform based on fit.

Anti-Pattern 4: Assuming the same model behaves and costs the same on every cloud

Treating a model on a partner cloud as identical to the first-party API: same price, same features, same quota, same release day. It is the same model family, not the same service.

Fix: Verify pricing, feature parity, regions, quotas, and data terms per cloud (§31.5), and run your contract tests against each deployment target (Module 37).


EXERCISE — Platform Mapping: Take a specific AI feature you are designing or maintaining. Map every component of its architecture to both a vendor-neutral implementation (using tools from the course) and a provider-native implementation (using your organization's primary cloud provider). For each component: note the trade-offs. Which components benefit most from the provider-native implementation? Which are better served by vendor-neutral approaches?

PONDER — The Lock-In Question: For any AI system your organization has built on a cloud provider's native services: if you needed to migrate to a different provider in 12 months, which components would be easiest to migrate? Which would be hardest? What would it cost in engineering time? Is the current lock-in level acceptable given your organization's strategy? Consider a concrete test: if the service you built on were renamed, frozen, or retired next year, as happened to Bedrock Agents and Prompt Flow in 2026, what would you have to rewrite?

WORKSHOP — AgentCore vs. DIY Decision: Your organization is deploying a multi-step research agent that needs to run for up to 4 hours, access 15 internal Lambda-based tools, maintain memory across sessions, and integrate with your existing CloudWatch observability stack. Build the architecture comparison: (1) using AgentCore Runtime + Gateway + Memory + Observability vs. (2) a self-managed approach using LangGraph + custom session management + Weaviate memory store + LangSmith observability. Compare: engineering effort, operational overhead, cost at 1,000 agent tasks/day, and lock-in risk. Then add a third column: the same agent on a second cloud's managed runtime. What would you have to change? Make the call.


Next: Module 32 — AI Agent Implementation Learnings