- Added session management for stateful execution in `sessions.py`, including session creation, state documentation, and cleanup of idle sessions. - Introduced `SecureContainerManager` in `containers.py` for managing container lifecycle with security enforcement, including creation, starting, stopping, and removal of containers. - Updated server initialization to use the new session manager. - Enhanced container configuration to skip resource limits for very high values, indicating no enforcement. - Improved logging capabilities for container operations in the audit logger. - Refactored Jupyter backend to integrate with the new session management and resource limits handling.
36 KiB
MCP-Forge Architecture Document
Overview
MCP-Forge is an MCP (Model Context Protocol) server that provides intelligent code execution capabilities for AI agents. It enables agents to execute Python code in isolated containers while having access to other MCP tools, optimizing data processing by keeping large datasets in the execution environment rather than passing them through the LLM context.
Core Concept
Problem: When AI agents use MCP tools that return large amounts of data, the data must flow through the LLM context window, which is:
- Token-expensive
- Slow
- Limited by context window size
Solution: MCP-Forge provides a code execution environment where:
- MCP tools are available as Python functions
- Agents can write code to process data locally
- Only relevant results flow back to the agent
- The agent already knows tool signatures from its own configuration
Architecture
AI Agent (e.g., GitHub Copilot, Claude)
│
├─ MCP Protocol
│
MCP-Forge Server
│
├─ MCP Resources (Discovery)
│ ├─ tools/available
│ ├─ sessions/{id}/state
│ ├─ sessions/{id}/variables
│ └─ environments/list
│
├─ MCP Tools (Execution)
│ ├─ execute_python()
│ ├─ document_state()
│ └─ build_custom_environment()
│
├─ Execution Backends
│ ├─ Simple Backend (stateless)
│ └─ Jupyter Backend (stateful)
│
├─ Custom Environment Builder
│ ├─ UV Package Installer
│ ├─ Build Cache Manager
│ ├─ Security Validator
│ └─ Template Library
│
└─ Podman API (restricted)
│
├─ Pre-built Images
│ ├─ mcp-forge/python:3.11
│ ├─ mcp-forge/python:3.12
│ └─ mcp-forge/jupyter:latest
│
├─ Custom User Images
│ └─ mcp-forge/custom:{user-env-name}
│
└─ Isolated Execution Containers
├─ Python Runtime
├─ MCP Client (injected tools)
└─ Session Volumes
Components
1. MCP Server Interface
Resources
mcp://forge/tools/available
- Returns: List of MCP tool names available in execution environment
- Purpose: Agent discovers what tools it can use in generated code
- Example:
["github_search_repos", "filesystem_read", "sqlite_query"]
mcp://forge/tools/{tool_name}/schema (optional)
- Returns: Full JSON schema for a specific tool
- Purpose: Fallback if agent doesn't have tool definition
- Note: Typically not needed as agent has tools in its own config
mcp://forge/sessions/{session_id}/state
- Returns: Documented state for a session
- Structure:
{ "documented_variables": { "df": "Customer data, 1000 rows, columns: id, name, purchase_date, amount", "model": "Trained RandomForest, accuracy 0.87" }, "note": "Preprocessing complete, ready for analysis", "last_updated": "2026-02-06T10:30:00Z", "all_variables": ["df", "model", "temp", "result"] }
mcp://forge/sessions/{session_id}/variables
- Returns: List of all variables in session namespace (kernel introspection)
- Purpose: Quick check of what exists in session
mcp://forge/environment/info (optional)
- Returns: Python version, installed packages, system info
- Purpose: Agent can verify environment capabilities
mcp://forge/environments/list
- Returns: List of available custom environments and templates
- Structure:
{ "custom": [ { "name": "my-ml-env", "image": "mcp-forge/custom:my-ml-env", "created": "2026-02-06T10:00:00Z", "packages": ["numpy==1.24.0", "pandas==2.0.0"], "size_mb": 1234 } ], "templates": [ { "name": "ml-basic", "description": "Basic ML stack", "packages": ["numpy", "pandas", "scikit-learn"] }, { "name": "data-science", "description": "Data science stack", "packages": ["numpy", "pandas", "matplotlib", "seaborn"] } ] }
mcp://forge/audit/operations (optional)
- Returns: Audit log of container operations
- Purpose: Security monitoring, debugging
Tools
execute_python
Execute Python code in an isolated container with MCP tools available.
Parameters:
{
"code": "string (required) - Python code to execute",
"mcp_tools": "array (optional) - List of MCP tool names to inject",
"session_id": "string (optional) - Session ID for stateful execution (null = stateless)",
"backend": "string (optional) - 'simple' (default) or 'jupyter'",
"timeout": "integer (optional) - Max execution time in seconds (default: 300)",
"volumes": "object (optional) - Volume mount configuration",
"custom_image": "string (optional) - Custom environment name to use",
"environment": "string (optional) - Template environment name (e.g., 'ml-basic', 'data-science')"
}
Returns:
{
"success": true,
"stdout": "string - Standard output",
"stderr": "string - Standard error",
"result": "any - Return value of last expression",
"available_tools": ["list of tools that were injected"],
"execution_time": 1.23,
"session_id": "abc123 (if stateful)"
}
Example:
# Agent calls:
execute_python(
code="""
repos = github_search_repos(query="MCP servers", max_results=100)
high_quality = [r for r in repos['items'] if r['stars'] > 50]
print(f"Found {len(high_quality)} high-quality repos")
high_quality[:10] # Return top 10
""",
mcp_tools=["github_search_repos"],
session_id=None # stateless
)
document_state
Document important variables in a stateful session for later retrieval.
Parameters:
{
"session_id": "string (required) - Session to document",
"variables": "object (required) - Variable name -> description mapping",
"note": "string (optional) - General note about session state",
"clear": "boolean (optional) - Clear existing documentation (default: false)"
}
Returns:
{
"success": true,
"documented_count": 3,
"session_id": "abc123"
}
Example:
document_state(
session_id="abc123",
variables={
"df": "Customer purchase data, 1000 rows, preprocessed and cleaned",
"model": "Trained RandomForest classifier, 87% accuracy on test set"
},
note="Ready for prediction phase"
)
build_custom_environment
Build a custom container image with specified Python packages. This is the only way to install additional packages - pip install is not allowed during code execution for security reasons.
Parameters:
{
"name": "string (required) - Name for the custom environment (alphanumeric + hyphens)",
"base_image": "string (optional) - Base image to build from (default: 'python:3.11')",
"packages": "array (required) - List of Python package specifications",
"python_version": "string (optional) - Python version: '3.11', '3.12' (default: '3.11')",
"description": "string (optional) - Description of this environment"
}
Returns:
{
"success": true,
"image_name": "mcp-forge/custom:my-ml-env",
"image_id": "sha256:abc123...",
"build_time": 45.2,
"installed_packages": ["numpy==1.24.0", "pandas==2.0.0", "scikit-learn==1.3.0"],
"cache_hit": false
}
Example:
# Agent builds custom environment for ML work
build_custom_environment(
name="my-ml-env",
packages=[
"numpy>=1.24.0",
"pandas>=2.0.0",
"scikit-learn>=1.3.0",
"matplotlib>=3.7.0",
"seaborn>=0.12.0"
],
description="Machine learning environment with common libraries"
)
# Later, use the custom environment
execute_python(
code="import pandas as pd; df = pd.read_csv('data.csv'); ...",
custom_image="mcp-forge/custom:my-ml-env",
session_id="ml-analysis"
)
2. Execution Backends
Simple Backend (Default)
Purpose: Stateless, single-shot code execution
Implementation:
- Spawn Podman container with Python image
- Inject code and MCP tool functions into namespace
- Execute via
python -c - Capture stdout, stderr, result
- Destroy container
Characteristics:
- Fast startup
- No state between calls
- Ideal for data transformation tasks
- Lower memory footprint
Use cases:
- Data processing and filtering
- One-off computations
- Format conversions
- Quick analyses
Jupyter Backend
Purpose: Stateful, multi-step workflows
Implementation:
- Spawn Podman container with IPython kernel
- Keep kernel running for session lifetime
- Execute code cells via Jupyter protocol (ZMQ)
- Maintain namespace between executions
- Support rich output formats
Characteristics:
- State persists between calls
- Variable persistence
- Interactive workflow support
- Higher resource usage
Use cases:
- Multi-step data analysis
- Iterative model training
- Building up complex state
- Exploratory workflows
Session Management:
- Sessions identified by unique ID
- Automatic timeout after inactivity (configurable, default: 1 hour)
- Manual cleanup via session deletion
- Resource limits per session
3. Custom Environment Building
Design Philosophy
Security Principle: pip install is NOT allowed during code execution. All package installations must happen during a controlled build process.
Why:
- Prevents malicious package installation during execution
- Ensures reproducible environments
- Allows security scanning of dependencies
- Enables caching for performance
- Provides audit trail of what's installed
Build System: UV-Based
Technology Choice: UV - ultra-fast Python package installer
Benefits:
- 10-100x faster than pip
- Built-in caching
- Deterministic dependency resolution
- Lockfile support for reproducibility
- Compatible with pip package specifications
Implementation
Build Process:
def build_custom_environment(name, packages, base_image="python:3.11"):
"""
Build a custom container image with specified packages using UV.
Steps:
1. Generate Containerfile with UV installation
2. Create requirements specification
3. Build image with layer caching
4. Validate and tag image
5. Store metadata for reuse
"""
# Generate Containerfile
containerfile = f"""
FROM {base_image}
# Install UV (cached layer)
RUN curl -LsSf https://astral.sh/uv/install.sh | sh
ENV PATH="/root/.cargo/bin:$PATH"
# Non-root user
RUN useradd -m -u 1000 forge
USER forge
# Copy requirements (cache-friendly)
COPY requirements.txt /tmp/requirements.txt
# Install packages with UV (uses cache)
RUN uv pip install --system -r /tmp/requirements.txt
# MCP bridge client
COPY mcp_tools.py /usr/local/lib/python3.11/site-packages/
WORKDIR /workspace
"""
# Build with Podman
build_context = create_build_context(containerfile, packages)
image_id = podman.images.build(
path=build_context,
tag=f"mcp-forge/custom:{name}",
cache_from=["mcp-forge/cache:uv-base"], # UV cache layer
buildargs={
"BUILDKIT_INLINE_CACHE": "1"
}
)
return image_id
Package Cache Strategy
Multi-Layer Caching:
Layer 1: Base Python image (rarely changes)
└─ python:3.11-slim
Layer 2: UV installation (stable)
└─ UV binary + dependencies
Layer 3: Common packages (pre-cached)
└─ numpy, pandas, requests, etc.
Layer 4: User packages (specific to build)
└─ Custom requirements
UV Cache Directory:
- Mounted from host:
/var/cache/mcp-forge/uv - Persisted between builds
- Dramatically speeds up builds with similar dependencies
Cache Invalidation:
- Requirements hash changes → rebuild user layer only
- Base image updated → rebuild from base
- UV version updated → rebuild from UV layer
Pre-built Environment Templates
Common Templates:
templates:
ml-basic:
description: "Basic ML stack"
packages:
- "numpy>=1.24.0"
- "pandas>=2.0.0"
- "scikit-learn>=1.3.0"
ml-advanced:
description: "Advanced ML with deep learning"
packages:
- "numpy>=1.24.0"
- "pandas>=2.0.0"
- "scikit-learn>=1.3.0"
- "torch>=2.0.0"
- "transformers>=4.30.0"
data-science:
description: "Data science stack"
packages:
- "numpy>=1.24.0"
- "pandas>=2.0.0"
- "matplotlib>=3.7.0"
- "seaborn>=0.12.0"
- "plotly>=5.14.0"
- "jupyter>=1.0.0"
web-scraping:
description: "Web scraping tools"
packages:
- "requests>=2.31.0"
- "beautifulsoup4>=4.12.0"
- "lxml>=4.9.0"
- "selenium>=4.10.0"
Usage:
# Use a template
build_custom_environment(
name="my-ml-env",
template="ml-basic",
additional_packages=["xgboost>=2.0.0"]
)
# Or build from scratch
build_custom_environment(
name="custom-env",
packages=["specific-package==1.0.0"]
)
Build Validation
Post-Build Checks:
def validate_built_environment(image_id):
"""
Validate that the built environment is safe and functional.
Checks:
1. All requested packages are installed
2. No malicious packages (against allowlist/blocklist)
3. Image size is within limits
4. Security scan passes (trivy/grype)
5. Python imports work
"""
# Import test
test_code = """
import sys
import json
installed = [(pkg.key, pkg.version) for pkg in __import__('pkg_resources').working_set]
print(json.dumps(installed))
"""
result = run_in_container(image_id, test_code)
installed_packages = json.loads(result.stdout)
# Security scan
scan_result = security_scan_image(image_id)
if scan_result.critical_vulns > 0:
raise SecurityError(f"Image has {scan_result.critical_vulns} critical vulnerabilities")
return {
"installed_packages": installed_packages,
"security_scan": scan_result,
"valid": True
}
Package Allowlist/Blocklist
Security Control:
packages:
# Automatically allowed (common, vetted)
allowlist:
- numpy
- pandas
- scikit-learn
- matplotlib
- seaborn
- requests
- beautifulsoup4
# ... many more
# Explicitly forbidden
blocklist:
- "os-sys" # Known malicious
- "cryptography-backdoor"
# Packages with known vulnerabilities
# Requires manual approval
manual_approval:
- "*crypto*" # Cryptographic packages
- "*network*" # Network access packages
- "*subprocess*" # Process spawning
Validation:
def validate_package_list(packages):
"""Check packages against allowlist/blocklist before building."""
for pkg in packages:
pkg_name = pkg.split("==")[0].split(">=")[0].split("<=")[0]
if pkg_name in BLOCKLIST:
raise SecurityError(f"Package {pkg_name} is blocked")
if not in_allowlist(pkg_name) and requires_approval(pkg_name):
raise ApprovalRequiredError(f"Package {pkg_name} requires manual approval")
Build Resource Limits
Prevent Build Abuse:
BUILD_LIMITS = {
"max_packages": 50, # Max packages per build
"max_build_time": 600, # 10 minutes
"max_image_size": "2GB", # Final image size
"max_concurrent_builds": 3, # Per user
"build_rate_limit": {
"requests": 10,
"period": 3600 # 10 builds per hour
}
}
Environment Lifecycle
Management:
# List user's custom environments
list_custom_environments() → [
{
"name": "my-ml-env",
"image": "mcp-forge/custom:my-ml-env",
"created": "2026-02-06T10:00:00Z",
"size_mb": 1234,
"packages": ["numpy==1.24.0", "pandas==2.0.0"],
"last_used": "2026-02-06T11:30:00Z"
}
]
# Delete unused environment
delete_custom_environment(name="old-env")
# Rebuild environment (e.g., after base image update)
rebuild_custom_environment(name="my-ml-env")
Auto-cleanup:
- Environments unused for 30 days → archived
- Archived for 90 days → deleted
- User notification before deletion
execute_python Integration
Using Custom Environments:
# Updated execute_python parameters
{
"code": "...",
"custom_image": "string (optional) - Custom environment name or image tag",
"environment": "string (optional) - Template name (alternative to custom_image)",
...
}
# Examples:
execute_python(
code="import numpy as np; ...",
custom_image="my-ml-env" # Use user's custom environment
)
execute_python(
code="import pandas as pd; ...",
environment="data-science" # Use pre-built template
)
4. Container Runtime (Podman)
Why Podman?
- Rootless by design: No daemon running as root
- Docker-compatible API: Easy integration
- Better security defaults: No privileged operations needed
- Daemonless: Containers are child processes
Deployment Modes
Development:
# Run mcp-forge directly on host
mcp-forge serve --podman-socket /run/user/1000/podman/podman.sock
Production:
# Run mcp-forge in container with socket mount
services:
mcp-forge:
image: mcp-forge:latest
volumes:
- /run/user/1000/podman/podman.sock:/run/podman/podman.sock:ro
environment:
- PODMAN_SOCKET=/run/podman/podman.sock
Security: Allowlist-Based Container Operations
Allowed Operations:
ALLOWED_OPERATIONS = {
"container.create": {
"allowed_images": [
"mcp-forge/python:3.11",
"mcp-forge/python:3.12",
"mcp-forge/jupyter:latest"
],
"forbidden_params": [
"privileged",
"cap_add",
"devices",
"pid_mode",
"ipc_mode"
],
"required_params": {
"network_mode": "none", # Or restricted network
"read_only": True, # Filesystem read-only except volumes
"memory_limit": "512m",
"cpu_quota": 50000, # 50% of one CPU
"security_opt": ["no-new-privileges"],
"user": "1000:1000" # Non-root user
}
},
"container.start": {
"session_containers_only": True # Only containers we created
},
"container.stop": {
"session_containers_only": True
},
"container.remove": {
"session_containers_only": True
},
"container.logs": {
"session_containers_only": True
}
}
Forbidden Operations:
FORBIDDEN_OPERATIONS = [
"container.exec", # No direct shell access
# Note: image.build IS allowed, but ONLY through build_custom_environment tool
# with validation, security scanning, and package allowlists
"image.pull", # No arbitrary image pulling (only pre-approved images)
"volume.create", # Only pre-configured volumes
"network.create", # No custom networks
"system.prune", # No system-level operations
]
# Additionally forbidden within execution containers:
EXECUTION_RESTRICTIONS = [
"subprocess.run(['pip', 'install', ...])", # No pip install during execution
"subprocess.run(['apt', 'install', ...])", # No system package installs
"import os; os.system('...')", # Restricted system calls
]
Volume Management
Allowed Volume Patterns:
ALLOWED_VOLUME_PATTERNS = [
"/mcp-forge/sessions/{session_id}/*", # Per-session data
"/mcp-forge/shared/readonly/*", # Shared read-only data
"/mcp-forge/uploads/{session_id}/*" # User uploads
]
FORBIDDEN_MOUNT_PATHS = [
"/",
"/etc",
"/var/run/docker.sock",
"/var/run/podman/podman.sock",
"/sys",
"/proc"
]
Resource Limits (Enforced)
RESOURCE_LIMITS = {
"memory": {
"default": "512m",
"max": "2g"
},
"cpu_quota": {
"default": 50000, # 50% of one CPU
"max": 100000 # 100% of one CPU
},
"timeout": {
"default": 300, # 5 minutes
"max": 1800 # 30 minutes
},
"pids_limit": 100,
"storage": {
"default": "1g",
"max": "10g"
}
}
5. MCP Tool Injection
Tool Discovery Flow
- Agent reads resource:
mcp://forge/tools/available - MCP-Forge returns:
["github_search_repos", "filesystem_read"] - Agent knows signatures: Already configured in agent's tool config
- Agent generates code: Using these tool names as Python functions
Tool Injection Implementation
At container startup:
# MCP-Forge generates injection code
def inject_mcp_tools(container, mcp_client, tool_names):
"""
Inject MCP tools as Python functions into container namespace.
Args:
container: Podman container instance
mcp_client: Connected MCP client with available tools
tool_names: List of tool names to inject
"""
injection_code = """
# MCP Tool Functions (auto-injected)
import json
from typing import Any
def _mcp_call(tool_name: str, **kwargs) -> Any:
'''Internal: Call MCP tool via bridge'''
# This communicates with MCP-Forge server which forwards to MCP client
import socket
s = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
s.connect('/tmp/mcp-bridge.sock')
s.send(json.dumps({'tool': tool_name, 'params': kwargs}).encode())
result = json.loads(s.recv(1024*1024).decode())
s.close()
return result
"""
# Generate wrapper function for each tool
for tool_name in tool_names:
tool_schema = mcp_client.get_tool_schema(tool_name)
# Generate function signature from schema
params = tool_schema.get('inputSchema', {}).get('properties', {})
param_str = ', '.join(f"{k}: Any = None" for k in params.keys())
injection_code += f"""
def {tool_name}({param_str}):
'''
{tool_schema.get('description', '')}
Auto-generated wrapper for MCP tool: {tool_name}
'''
kwargs = {{{', '.join(f"'{k}': {k}" for k in params.keys())}}}
kwargs = {{k: v for k, v in kwargs.items() if v is not None}}
return _mcp_call('{tool_name}', **kwargs)
"""
# Write to container's Python site-packages or PYTHONSTARTUP
container.write_file('/usr/local/lib/python3.11/site-packages/mcp_tools.py',
injection_code)
# Start MCP bridge server (forwards calls to actual MCP client)
start_mcp_bridge(container, mcp_client)
Alternative: Environment Variable Approach
# Set PYTHONSTARTUP to auto-import tools
container.env['PYTHONSTARTUP'] = '/mcp-forge/startup.py'
MCP Bridge Server
A lightweight server running in the MCP-Forge process that:
- Listens on Unix socket mounted into container
- Receives tool calls from container code
- Forwards to actual MCP client
- Returns results to container
This keeps the MCP client logic outside the untrusted container.
6. Session State Management
State Lifecycle
Stateless Execution (session_id=None):
Request → Create Container → Execute → Return Results → Destroy Container
Stateful Execution (session_id="abc123"):
First Request:
Request → Create Container → Start Kernel → Execute → Return Results
Container stays alive
Subsequent Requests:
Request → Reuse Container → Execute in Same Kernel → Return Results
Cleanup:
Timeout or Manual → Stop Container → Destroy Container
State Documentation
Agent Workflow:
# Step 1: Load and process data
execute_python(
code="df = load_large_dataset(); df_clean = preprocess(df)",
session_id="analysis-123",
backend="jupyter"
)
# Step 2: Document important state
document_state(
session_id="analysis-123",
variables={
"df": "Raw dataset, 100k rows",
"df_clean": "Cleaned dataset, 95k rows, ready for modeling"
},
note="Data loading complete"
)
# ... conversation continues, context window rotates ...
# Step 3: Agent checks state after context rotation
state = read_resource("mcp://forge/sessions/analysis-123/state")
# Agent sees: df_clean exists and is ready
# Step 4: Continue work
execute_python(
code="model = train_model(df_clean); results = evaluate(model)",
session_id="analysis-123"
)
State Introspection
Automatic tracking:
- Kernel namespace inspection (all variables)
- Variable types and sizes
- Last execution timestamp
Agent-documented:
- Semantic descriptions
- Workflow notes
- Important variables highlighted
Combined in resource:
{
"all_variables": ["df", "df_clean", "model", "temp", "i", "results"],
"documented_variables": {
"df_clean": "Cleaned dataset, 95k rows, ready for modeling",
"model": "Trained model, 87% accuracy",
"results": "Evaluation metrics"
},
"introspection": {
"df_clean": {"type": "DataFrame", "shape": [95000, 12], "memory_mb": 87},
"model": {"type": "RandomForestClassifier", "memory_mb": 234}
},
"note": "Ready for final predictions",
"last_updated": "2026-02-06T10:30:00Z"
}
Implementation Considerations
Technology Stack
MCP Server:
- Language: Python 3.11+
- Framework:
mcpSDK (official Python implementation) - Async:
asynciofor concurrent operations
Container Runtime:
- Podman via
podman-pylibrary - Fallback: Direct Podman CLI calls
Jupyter Backend:
jupyter_clientfor kernel management- ZMQ for communication
ipykernelin container images
Simple Backend:
- Direct Python execution via
subprocess - Or:
python:3.11-slimbase image withexecentrypoint
Pre-built Container Images
Base Image: mcp-forge/python:3.11
FROM python:3.11-slim
# Non-root user
RUN useradd -m -u 1000 forge
USER forge
# Install common packages
RUN pip install --user numpy pandas requests
# MCP bridge client
COPY mcp_tools.py /usr/local/lib/python3.11/site-packages/
# Startup script
COPY startup.py /mcp-forge/startup.py
WORKDIR /workspace
Jupyter Image: mcp-forge/jupyter:latest
FROM mcp-forge/python:3.11
USER root
RUN pip install ipykernel jupyter_client
USER forge
# IPython config for security
COPY ipython_config.py /home/forge/.ipython/profile_default/
CMD ["python", "-m", "ipykernel_launcher", "-f", "/tmp/kernel.json"]
Error Handling
Container Failures:
- Timeout: Kill container, return timeout error
- OOM: Return memory limit error with suggestion to increase
- Crash: Return stderr and exit code
MCP Tool Failures:
- Tool not available: Clear error message with available tools
- Tool call error: Return tool error to agent, don't crash execution
- Network error: Retry logic for transient failures
Security Violations:
- Forbidden operation: Reject immediately, log to audit
- Resource limit exceeded: Terminate execution, clear error message
- Invalid volume mount: Reject with explanation
Performance Considerations
Container Reuse:
- Pool of warm containers for simple backend (optional)
- Lazy cleanup of idle Jupyter sessions
Parallel Execution:
- Support multiple concurrent executions
- Per-session locking for stateful operations
- Configurable max concurrent containers
Image Caching:
- Pre-pull images on startup
- Periodic image updates (configurable schedule)
Volume Performance:
- Use tmpfs for ephemeral data
- Persistent volumes for session data
- Cleanup strategy for old sessions
Security Model
Defense in Depth
- Podman rootless: No root daemon
- Allowlist enforcement: Only permitted operations
- Read-only filesystem: Except specific volumes
- Network isolation: No internet by default
- Resource limits: CPU, memory, storage, PIDs
- No privileged mode: Ever
- Capability dropping: Minimal capabilities
- User namespaces: Non-root user in container
- Audit logging: All operations logged
- Session isolation: Each session has isolated volumes
Threat Model
Threats Mitigated:
- Container escape → Rootless + restricted operations
- Resource exhaustion → Hard limits enforced
- Data exfiltration → Network isolation
- Privilege escalation → No privileged mode, capability restrictions
- Host filesystem access → Allowlist volumes only
Out of Scope:
- Side-channel attacks (Spectre, Meltdown)
- Physical security
- Supply chain attacks on base images (use trusted registries)
Audit and Monitoring
Logged Events:
- Container create/start/stop/remove
- Execution requests (code hash, not full code for privacy)
- Resource limit violations
- Security policy violations
- Session lifecycle events
Log Format:
{
"timestamp": "2026-02-06T10:30:00Z",
"event": "container.create",
"session_id": "abc123",
"image": "mcp-forge/python:3.11",
"resources": {"memory": "512m", "cpu_quota": 50000},
"success": true
}
Configuration
Server Configuration
Example mcp-forge.yaml:
server:
host: localhost
port: 3000
podman_socket: /run/user/1000/podman/podman.sock
execution:
default_backend: simple
default_timeout: 300
max_timeout: 1800
default_memory: 512m
max_memory: 2g
default_cpu_quota: 50000
max_cpu_quota: 100000
images:
python_3_11: mcp-forge/python:3.11
python_3_12: mcp-forge/python:3.12
jupyter: mcp-forge/jupyter:latest
auto_pull: true
pull_interval: 86400 # 24 hours
sessions:
idle_timeout: 3600 # 1 hour
max_concurrent: 10
cleanup_interval: 300 # 5 minutes
volumes:
base_path: /var/lib/mcp-forge
session_quota: 1g
max_session_quota: 10g
security:
audit_log: /var/log/mcp-forge/audit.log
enforce_resource_limits: true
allow_network: false
environment_builder:
enabled: true
uv_cache_path: /var/cache/mcp-forge/uv
max_packages_per_build: 50
max_build_time: 600 # 10 minutes
max_image_size: 2147483648 # 2GB in bytes
max_concurrent_builds: 3
build_rate_limit:
requests: 10
period: 3600 # 10 builds per hour per user
# Package security
package_validation:
use_allowlist: true
allowlist_path: /etc/mcp-forge/package-allowlist.txt
blocklist_path: /etc/mcp-forge/package-blocklist.txt
require_approval_patterns:
- "*crypto*"
- "*network*"
- "*subprocess*"
# Environment lifecycle
auto_cleanup:
enabled: true
archive_after_days: 30
delete_after_days: 90
# Pre-built templates
templates:
ml-basic:
description: "Basic ML stack"
packages:
- "numpy>=1.24.0"
- "pandas>=2.0.0"
- "scikit-learn>=1.3.0"
data-science:
description: "Data science stack"
packages:
- "numpy>=1.24.0"
- "pandas>=2.0.0"
- "matplotlib>=3.7.0"
- "seaborn>=0.12.0"
- "plotly>=5.14.0"
web-scraping:
description: "Web scraping tools"
packages:
- "requests>=2.31.0"
- "beautifulsoup4>=4.12.0"
- "lxml>=4.9.0"
mcp_tools:
# MCP servers to connect to and expose
github:
command: "npx"
args: ["-y", "@modelcontextprotocol/server-github"]
env:
GITHUB_TOKEN: ${GITHUB_TOKEN}
filesystem:
command: "npx"
args: ["-y", "@modelcontextprotocol/server-filesystem", "/allowed/path"]
Environment Variables
# Podman socket location
MCP_FORGE_PODMAN_SOCKET=/run/user/1000/podman/podman.sock
# Base path for volumes
MCP_FORGE_VOLUMES_PATH=/var/lib/mcp-forge
# UV cache for fast package installation
MCP_FORGE_UV_CACHE_PATH=/var/cache/mcp-forge/uv
# Security
MCP_FORGE_AUDIT_LOG=/var/log/mcp-forge/audit.log
# MCP tool credentials
GITHUB_TOKEN=ghp_xxxxx
API Examples
Example 1: Simple Data Processing
# Agent discovers available tools
tools = read_resource("mcp://forge/tools/available")
# Returns: ["github_search_repos", "filesystem_read"]
# Agent generates and executes code
result = execute_python(
code="""
repos = github_search_repos(query="machine learning", max_results=100)
python_repos = [r for r in repos['items'] if r['language'] == 'Python']
top_10 = sorted(python_repos, key=lambda x: x['stars'], reverse=True)[:10]
# Return summary
[{
'name': r['name'],
'stars': r['stars'],
'url': r['html_url']
} for r in top_10]
""",
mcp_tools=["github_search_repos"],
session_id=None
)
print(result['result'])
# Returns top 10 Python ML repos
Example 2: Building and Using Custom Environment
# Agent needs ML libraries for analysis
# First, check available environments
envs = read_resource("mcp://forge/environments/list")
# Build custom environment if needed
build_result = build_custom_environment(
name="my-ml-analysis",
packages=[
"numpy>=1.24.0",
"pandas>=2.0.0",
"scikit-learn>=1.3.0",
"matplotlib>=3.7.0",
"xgboost>=2.0.0"
],
description="Custom ML environment for analysis project"
)
# Wait for build to complete (typically 30-60 seconds with UV cache)
print(f"Built {build_result['image_name']} in {build_result['build_time']}s")
print(f"Installed: {build_result['installed_packages']}")
# Now use the custom environment for analysis
result = execute_python(
code="""
import pandas as pd
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
import xgboost as xgb
# Load data (from volume or MCP tool)
data = filesystem_read('/data/customer_churn.csv')
df = pd.read_csv(data)
# Preprocess
X = df.drop('churn', axis=1)
y = df['churn']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)
# Train models
rf_model = RandomForestClassifier(n_estimators=100)
rf_model.fit(X_train, y_train)
rf_score = rf_model.score(X_test, y_test)
xgb_model = xgb.XGBClassifier()
xgb_model.fit(X_train, y_train)
xgb_score = xgb_model.score(X_test, y_test)
# Return results
{
'random_forest_accuracy': rf_score,
'xgboost_accuracy': xgb_score,
'best_model': 'xgboost' if xgb_score > rf_score else 'random_forest'
}
""",
custom_image="my-ml-analysis",
mcp_tools=["filesystem_read"],
session_id="churn-analysis"
)
print(result['result'])
# Returns: {'random_forest_accuracy': 0.87, 'xgboost_accuracy': 0.91, 'best_model': 'xgboost'}
Example 3: Stateful Analysis with Custom Environment
# Create session with custom environment
session_id = "deep-analysis"
# Step 1: Load and explore data
execute_python(
code="""
import pandas as pd
import matplotlib.pyplot as plt
df = filesystem_read_csv('/data/sales_2024.csv')
print(f"Loaded {len(df)} rows")
print(df.describe())
# Store for later
df.to_pickle('/tmp/sales_df.pkl')
""",
custom_image="data-science", # Use pre-built template
mcp_tools=["filesystem_read_csv"],
session_id=session_id,
backend="jupyter"
)
# Document what we have
document_state(
session_id=session_id,
variables={
"df": "Sales data for 2024, 50k rows, pickled to /tmp/sales_df.pkl"
},
note="Data loaded and ready for analysis"
)
# ... conversation continues, context rotates ...
# Step 2: Later, continue analysis
state = read_resource(f"mcp://forge/sessions/{session_id}/state")
# Agent sees df is available
execute_python(
code="""
import pandas as pd
df = pd.read_pickle('/tmp/sales_df.pkl')
# Analyze trends
monthly_sales = df.groupby('month')['revenue'].sum()
growth_rate = monthly_sales.pct_change().mean() * 100
{
'total_revenue': float(df['revenue'].sum()),
'avg_monthly_growth': float(growth_rate),
'top_products': df.groupby('product')['revenue'].sum().nlargest(5).to_dict()
}
""",
session_id=session_id
)
Example 4: Using Pre-built Templates
# Quick start with pre-built template
result = execute_python(
code="""
import requests
from bs4 import BeautifulSoup
# Scrape some data
response = requests.get('https://example.com/data')
soup = BeautifulSoup(response.content, 'lxml')
# Extract and process
data = [item.text for item in soup.find_all('div', class_='data-item')]
data[:10] # Return first 10 items
""",
environment="web-scraping", # Use template instead of building custom
session_id=None
)
print(result['result'])
Implementation Roadmap
Phase 1: Core Execution (MVP)
- Basic execute_python tool with simple backend
- Podman integration with security restrictions
- MCP tool injection
- Pre-built Python 3.11 image
Phase 2: Stateful Execution
- Jupyter backend implementation
- Session management
- State documentation tool
- Session resources
Phase 3: Custom Environments
- build_custom_environment tool
- UV-based package installation
- Build caching
- Package allowlist/blocklist
- Security validation
Phase 4: Advanced Features
- Pre-built template library
- Environment lifecycle management
- Build rate limiting
- Advanced audit logging
- Multi-user support with quotas
Phase 5: Production Hardening
- Performance optimization
- Advanced security scanning
- Comprehensive monitoring
- High availability setup
- Documentation and examples