Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
Stop Building AI Apps for Every Idea. Start Building MCP Servers — Part #5
Artificial Intelligence   Latest   Machine Learning

Stop Building AI Apps for Every Idea. Start Building MCP Servers — Part #5

Author(s): Andrii Tkachuk

Originally published on Towards AI.

Stop Building AI Apps for Every Idea. Start Building MCP Servers — Part #5

Where does your MCP server actually live?

In Part 1 I argued that the center of gravity in applied AI is shifting — the UI is becoming the shell, and MCP servers are becoming the capability layer.

Stop Building AI Apps for Every Idea. Start Building MCP Servers — Part #1

For many teams, the playbook still looks the same:

ai.plainenglish.io

In Part 2 we looked at what it takes to survive production: transport choices, stateless HTTP, sampling, roots, OAuth, and deployment topology.

Stop Building AI Apps for Every Idea. Start Building MCP Servers — Part #2

Over the past year, the Model Context Protocol has quietly moved from “interesting Anthropic experiment” to the default…

pub.towardsai.net

In Part 3 we covered runtime context — why user_id, tenant_id, and other trusted identifiers must never travel through the model.

Stop Building AI Apps for Every Idea. Start Building MCP Servers — Part #3

In Part 1 we talked about why you should be building MCP servers instead of yet another AI app. In Part 2 — how to…

pub.towardsai.net

In Part 4 we went deep on security architecture: policy registries, tool filtering, execution gateways, and audit trails.

Stop Building AI Apps for Every Idea. Start Building MCP Servers — Part #4

In Part 1, I argued that the center of gravity in applied AI is shifting from full applications to MCP servers. The UI…

pub.towardsai.net

This part is about the one question nobody seems to answer clearly:

Once your MCP server is built, where does it actually run in production?

Stop Building AI Apps for Every Idea. Start Building MCP Servers — Part #5
Photo by Nathan Duck on Unsplash

Before we start! 🦾

If this piece gives you something practical you can take into your own system:
👏 leave 50 claps (yes, you can!) — Medium’s algorithm favors this, increasing visibility to others who then discover the article.
🔔 Follow me on Medium and LinkedIn for more deep dives into agentic systems, LLM architecture, and production-grade AI engineering.

Why deployment finally deserves its own article

The four previous parts covered what to build and how to build it well. But there’s a structural question sitting underneath all of it that shapes every other decision:

Where does the server live?

The answer matters more than it used to. MCP has crossed a threshold. It is no longer just a local developer tool wired to Claude Desktop over stdio. It is now a network protocol with production-grade semantics: authenticated remote endpoints, session management, streaming responses, multi-tenant isolation, and IAM-compatible auth flows. The deployment question has become a first-class architectural concern.

And the ecosystem has responded. In 2026, you have three meaningfully different deployment philosophies to choose from:

  1. Self-managed containers — you own the Docker image, you own the runtime (Lambda, ECS, Kubernetes, whatever your team already operates)
  2. Managed hosting — AWS AgentCore Runtime takes the infrastructure off your hands
  3. The Docker Hub MCP ecosystem — a new distribution model that changes how MCP servers are discovered, packaged, and consumed

Each of these is a real option. Each has its own tradeoffs. And the good news — which I’ll say up front — is that MCP is now fully serverless-compatible. The protocol’s support for streamable-http with stateless_http=True means Lambda has been an open door for a while. If you're building on Docker images, you're not locked to any specific runtime. That's the real shift.

The container-first principle

Before going into the specific options, there’s one principle that should guide all of them:

Build once. Deploy anywhere. The image is the unit of portability.

Your MCP server is a FastMCP application running inside a container. That’s your artifact. What you do with that artifact — whether you push it to Lambda, ECS Fargate, AgentCore, Kubernetes, or a plain EC2 host — is an infrastructure decision that should be completely orthogonal to your application logic.

This matters because teams sometimes treat deployment targets as permanent choices. They’re not. If your server runs cleanly in a Docker image, you can move it between runtimes with nothing more than a config change and a redeploy. That portability is worth protecting — which means never coupling your server code to any specific deployment platform.

The way this plays out in FastMCP:

mcp = FastMCP(
name="your-server",
host="0.0.0.0",
port=8000,
stateless_http=True, # or False, depending on your protocol needs
)

if __name__ == "__main__":
mcp.run(transport="streamable-http")

That code works in a local Docker container, in a Lambda Web Adapter, in ECS Fargate, and in AgentCore Runtime. The entrypoint doesn’t change. The Dockerfile doesn’t change. The deployment target is a delivery detail, not an architectural one.

Option 1: Self-managed containers (Lambda / ECS / K8s)

This is the path most teams default to, and for good reason: it maps exactly to infrastructure your team already knows how to operate, monitor, and debug.

Lambda — the right fit for stateless, bursty tools

Lambda is now a legitimate deployment target for MCP servers. The combination of streamable-http transport and stateless_http=True means your server can run as a normal HTTP handler. Two deployment paths work:

Lambda Web Adapter (recommended): The AWS Lambda Web Adapter project lets you lift your FastMCP server directly onto Lambda without rewriting anything. It translates API Gateway events into the HTTP payload FastMCP expects on its socket, handles initialization and graceful shutdown, and manages the Lambda lifecycle transparently. Your code thinks it’s a normal HTTP server. Lambda handles the rest.

Native Lambda: You can also implement MCP message handling directly in the Lambda handler without running a full HTTP server inside the function. More control, more complexity — usually not worth it unless you have specific cold start constraints.

Lambda’s value proposition for MCP is clear: pay-per-use economics on bursty, agent-driven traffic. A single user prompt can trigger 5–50 tool calls in rapid succession. Lambda scales to absorb that burst automatically and bills per millisecond of execution. An idle server costs nothing.

The hard limits are equally clear. Lambda has a maximum execution time of 15 minutes per invocation. More critically: Lambda is stateless. This is fine for most tools, but it means you cannot use sampling, server-initiated requests, or meaningful progress streaming. If your MCP server’s value depends on any of those protocol features, Lambda is not the right target. It’s the right tool for narrow, fast, stateless capability endpoints — not for servers that need a bidirectional conversation with the client.

Lambda → Good for:
stateless tools
bursty / sporadic traffic
cost-sensitive deployments
internal developer tooling with erratic usage


Lambda → Not for:
sampling (server-initiated LLM calls)
elicitation (multi-turn tool execution)
meaningful progress streaming
sessions > 15 minutes

ECS Fargate — the workhorse for production servers

For servers that need session state, long-running tools, or the full MCP protocol feature set, ECS Fargate is the practical sweet spot. You push your Docker image, configure a task definition, and Fargate runs your container as a long-lived service without any underlying instance management.

The deployment story is simple: your FastMCP container exposes port 8000 at /mcp. An Application Load Balancer terminates TLS and forwards traffic to your tasks. ECS handles scaling, health checks, and task replacement.

For stateful servers (stateless_http=False), you need sticky sessions on the load balancer — ALB has this natively. Each MCP session maintains a persistent SSE channel, and all requests for that session need to reach the same container instance. Configure your target group with stickiness.enabled=true and an appropriate duration. This is a load balancer config, not an application change.

For stateless servers, no sticky sessions needed. Any task behind the ALB can handle any request. Horizontal scaling is trivial.

Cost perspective: ECS Fargate at typical MCP server scale — a few hundred requests per hour from internal tooling — runs comfortably under $10/month. One practitioner quoted $3/month for a production internal tools server. Lambda is cheaper at very low volume; Fargate is cheaper at consistent sustained load.

ECS Fargate → Good for:
full MCP protocol (sampling, elicitation, progress)
stateful sessions
consistent sustained traffic
warm caches and in-process state
any language / any runtime
teams already running containers on AWS

ECS Fargate → Trade-offs:
container lifecycle management (you own it)
sticky session config for stateful mode
slightly higher baseline cost than Lambda

Kubernetes — when you’re already there

If your team runs Kubernetes, there’s nothing special about MCP servers from the platform perspective. A FastMCP container is a standard HTTP workload. The MCP-specific considerations are exactly the same as ECS: stateless mode gets simple deployments, stateful mode needs session affinity (an ingress annotation, not a code change).

Become a Medium member

The reason to choose Kubernetes over ECS isn’t MCP — it’s your existing operational investment. If you have observability stacks, GitOps pipelines, and on-call runbooks built around Kubernetes, deploying MCP servers there costs almost nothing incrementally. Don’t migrate away from Kubernetes for MCP. Don’t migrate to Kubernetes for MCP either.

Option 2: AWS AgentCore Runtime — the managed path

AWS Bedrock AgentCore Runtime is purpose-built for hosting agents and MCP servers. It launched with stateless support and added stateful MCP capabilities in March 2026. It is now a first-class deployment target that deserves serious consideration, especially for AWS-native teams.

What AgentCore actually gives you

The headline feature is what AgentCore handles at the platform level rather than leaving it to your application:

Session isolation by microVM: In stateful mode, each MCP session runs in a completely isolated execution environment — not just a separate request, a separate microVM. This is meaningfully different from ECS stateful deployments where you’re managing session affinity yourself at the load balancer. AgentCore’s isolation is deeper and automatic.

Serverless + stateful simultaneously: This is the technical trick that AgentCore pulls off that Lambda cannot. Lambda is serverless but stateless. ECS is stateful but requires persistent container management. AgentCore gives you pay-per-processing-time economics and session state preservation across multiple requests within the same session. For many MCP server workloads, that’s the ideal combination.

Built-in IAM and OAuth/Cognito auth: Deploy to AgentCore and you get authentication handled by the platform. IAM SigV4 signing is supported out of the box. OAuth with Cognito requires a user pool setup but plugs in directly. You’re not writing auth middleware; you’re configuring the runtime.

Full MCP protocol including sampling and elicitation: Since March 2026, AgentCore supports stateful MCP — servers can initiate requests back to the client, run multi-turn elicitation flows to collect user input interactively during tool execution, invoke LLM sampling for dynamic content generation, and stream real-time progress notifications. The full protocol is available.

The deployment workflow

The AgentCore CLI makes the deploy path almost trivially simple:

# Install the CLI
npm install -g @aws/agentcore

# Scaffold a new project with MCP protocol
agentcore create --protocol MCP

# Copy your server code into the generated structure
# agentcore/agentcore.json points to your entrypoint

# Deploy
agentcore deploy

The CLI packages your code, uploads to S3, creates a runtime, and returns an ARN. Your server is now reachable at a Bedrock AgentCore invocation endpoint. The path convention is 0.0.0.0:8000/mcp — which is the FastMCP default. In most cases, you copy your existing server file in and deploy with no code changes.

The constraints to understand

AgentCore is not a universal answer. A few specific constraints matter:

It expects stateless HTTP by default. The runtime recommends stateless_http=True for basic servers, and automatically injects Mcp-Session-Id headers for continuity. Stateful mode is available but requires explicit configuration.

It’s AWS-specific. If your organization runs across multiple clouds or you want to avoid platform dependency, AgentCore locks you to AWS infrastructure and IAM. For some teams that’s fine; for others, it’s a non-starter.

The MCP server is not the same as the AgentCore MCP server (the tooling). AWS also ships a separate MCP server about AgentCore — a meta-tool for Claude Code and Cursor that helps you manage AgentCore deployments through conversational commands. Don’t confuse the two: one is your server deployed on AgentCore; the other is AWS tooling for managing AgentCore.

AgentCore Runtime → Good for:
AWS-native teams
serverless + stateful simultaneously
full MCP protocol (sampling, elicitation)
zero session management overhead
teams that want managed infrastructure end-to-end

AgentCore Runtime → Trade-offs:
AWS-only (no multi-cloud portability)
Cognito setup required for OAuth
less operational control than ECS
vendor lock-in at the runtime layer

Option 3: The Docker Hub MCP Ecosystem — a new distribution model

This one is different. It’s not just a deployment target — it’s a distribution and discovery layer that changes how MCP servers move from builders to consumers.

What Docker Hub MCP actually is

Docker launched their MCP Catalog as a curated registry of verified MCP servers packaged as container images. At the time of writing, it hosts 300+ servers. These are not arbitrary Docker images — they are purpose-built MCP servers with cryptographic signatures, Software Bill of Materials (SBOMs), provenance tracking, and automatic security updates maintained by Docker.

The catalog includes official servers from Stripe, Elastic, GitHub, New Relic, Neo4j, Grafana, Heroku, MongoDB, and many others. The Docker Hub MCP server itself is in the catalog — an MCP server that lets LLMs discover and query Docker Hub repositories through natural language.

From a builder’s perspective, the interesting part is publishing. If you build an MCP server and submit it to the Docker MCP Registry via pull request, Docker will build, sign, and publish your image to mcp/your-server-name on Docker Hub. Your server becomes discoverable by any AI client connected to the catalog.

Docker MCP Gateway: the local runtime

The Docker MCP Toolkit ships with an MCP Gateway that manages the lifecycle of catalog servers locally. When an AI client requests a tool, the Gateway identifies which server provides it, starts it as an isolated container (if not already running), and proxies the request. The isolation is real: each server runs with restricted privileges, controlled network access, and scoped resource limits.

This creates an interesting operational model for internal tooling: your team runs the MCP Gateway locally or on a shared internal host, configures a custom catalog pointing to your internal servers, and every developer in the team connects their AI client to the same gateway. Auth is handled once; catalog is controlled centrally; servers are containerized and isolated.

For enterprise scenarios, custom catalogs matter. Instead of exposing the full 300+ server Docker catalog to your team, you define exactly which servers are available — your internal servers plus approved third-party ones. When combined with DockerAI Governance (launched June 2026), you get centralized control over which MCP tools agents can call, what they can reach on the network, and which credentials they can use.

How this changes distribution

The important shift here isn’t just convenience. It’s that Docker Hub MCP creates a pull-based distribution model for MCP capability.

Previously, if you wanted to give your tools to another team or organization, you had to give them a URL, an auth token, and documentation. They had to stand up an HTTP client, manage credentials, and maintain the connection. With Docker Hub MCP, the process becomes closer to: here’s an image on Docker Hub, connect to the catalog, and your AI client can discover and use it immediately.

This is still maturing, but the direction is clear. MCP servers are becoming packages — distributable, versioned, signed, and independently deployable artifacts — not just endpoints you point clients at.

Docker Hub MCP → Good for:
distributing MCP servers to other teams / organizations
teams already invested in Docker workflows
local developer tooling with isolation guarantees
building servers you want to publish publicly
internal tool catalogs with centralized governance

Docker Hub MCP → Trade-offs:
primarily local/on-prem in the current model
not a full remote hosting solution by itself
catalog-based discovery is still evolving

How to choose

Here’s the decision framework I actually use:

Do you need stateful sessions (sampling, elicitation, progress)?
├── No → Lambda (stateless, bursty) or AgentCore stateless
└── Yes → AgentCore Runtime or ECS Fargate
├── AWS-native, want zero infra management → AgentCore
└── Need operational control / multi-cloud → ECS Fargate

Traffic pattern?
├── Sporadic / < 1M requests/month → Lambda
├── Consistent sustained load → ECS Fargate
└── Variable, agent-driven bursts → AgentCore or Lambda

Distribution model?
├── Internal team tooling → any of the above
├── Publish for others to consume → Docker Hub MCP Catalog
└── Enterprise governance over agent tools → Docker MCP Gateway

For most internal enterprise tooling that my teams build, the practical answer is: ECS Fargate for production servers that need the full protocol, Lambda for narrow stateless tools that only need fast execution, AgentCore when the team is AWS-native and wants managed infrastructure end-to-end.

The Docker Hub MCP angle is increasingly relevant for teams that build tooling for other teams — not as the primary runtime, but as the distribution layer on top of it.

The one thing that doesn’t change across all options

Regardless of where your server runs, the code-level discipline from the previous parts applies uniformly:

Runtime context (user_id, tenant_id) always comes from the transport layer — headers or IAM claims — not from LLM-controlled arguments. Tool schemas stay narrow. Authorization enforces at execution time, not just at tool discovery. Audit events are written before the response is returned.

None of that changes based on deployment target. The architecture principles are deployment-agnostic. The infrastructure decision is just the final layer of the stack.

The image is your artifact. The runtime is a config choice. The capability layer is what actually matters.

Final take

MCP server deployment in 2026 is no longer a single answer. It’s a decision space:

Lambda gives you serverless economics for stateless tools. ECS Fargate gives you the full protocol with operational control. AgentCore gives you managed infrastructure with serverless + stateful simultaneously. Docker Hub MCP gives you a distribution model.

All of them work. All of them run the same FastMCP image. The portability is real — protect it by keeping your application code decoupled from any specific runtime, and you can migrate between targets as your requirements evolve.

The series so far:

  • Part 1: Why MCP servers are replacing AI apps as the default unit of work
  • Part 2: Making them production-ready (transport, auth, sessions, templates)
  • Part 3: Runtime context — what the model must never touch
  • Part 4: Security architecture — authorization, policy registries, audit trails
  • Part 5 (this): Where it actually runs — Lambda, ECS, AgentCore, Docker Hub MCP

And that’s a wrap! If you’ve read this far, it probably means you found this article useful or insightful. If that’s the case, consider leaving a few claps or sharing it with your team, please. Thanks for reading! 🚀

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day

→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.