mcp-forge/docs/architecture2.md

1475 lines
37 KiB
Markdown

# MCP-Forge Architecture Document
## Overview
MCP-Forge is an MCP (Model Context Protocol) server that provides intelligent code execution capabilities for AI agents. It enables agents to execute Python and R code in isolated containers while having access to other MCP tools, optimizing data processing by keeping large datasets in the execution environment rather than passing them through the LLM context.
## Core Concept
**Problem:** When AI agents use MCP tools that return large amounts of data, the data must flow through the LLM context window, which is:
- Token-expensive
- Slow
- Limited by context window size
**Solution:** MCP-Forge provides a code execution environment where:
- MCP tools are available as language-native functions (Python and R)
- Agents can write code to process data locally
- Only relevant results flow back to the agent
- The agent already knows tool signatures from its own configuration
## Architecture
```
AI Agent (e.g., GitHub Copilot, Claude)
├─ MCP Protocol
MCP-Forge Server
├─ MCP Resources (Discovery)
│ ├─ tools/available
│ ├─ sessions/{id}/state
│ ├─ sessions/{id}/variables
│ └─ environments/list
├─ MCP Tools (Execution)
│ ├─ execute_python()
│ ├─ document_state()
│ └─ build_custom_environment()
├─ Execution Backends
│ ├─ Simple Backend (stateless)
│ └─ Jupyter Backend (stateful)
├─ Custom Environment Builder
│ ├─ UV Package Installer
│ ├─ Build Cache Manager
│ ├─ Security Validator
│ └─ Template Library
└─ Podman API (restricted)
├─ Pre-built Images
│ ├─ mcp-forge/python:3.11
│ ├─ mcp-forge/python:3.12
│ └─ mcp-forge/jupyter:latest
├─ Custom User Images
│ └─ mcp-forge/custom:{user-env-name}
└─ Isolated Execution Containers
├─ Python/R Runtime
├─ MCP Client (injected tools)
└─ Session Volumes
```
## Components
### 1. MCP Server Interface
#### Resources
**`mcp://forge/tools/available`**
- Returns: List of MCP tool names available in execution environment
- Purpose: Agent discovers what tools it can use in generated code
- Example: `["github_search_repos", "filesystem_read", "sqlite_query"]`
**`mcp://forge/tools/{tool_name}/schema`** (optional)
- Returns: Full JSON schema for a specific tool
- Purpose: Fallback if agent doesn't have tool definition
- Note: Typically not needed as agent has tools in its own config
**`mcp://forge/sessions/{session_id}/state`**
- Returns: Documented state for a session
- Structure:
```json
{
"documented_variables": {
"df": "Customer data, 1000 rows, columns: id, name, purchase_date, amount",
"model": "Trained RandomForest, accuracy 0.87"
},
"note": "Preprocessing complete, ready for analysis",
"last_updated": "2026-02-06T10:30:00Z",
"all_variables": ["df", "model", "temp", "result"]
}
```
**`mcp://forge/sessions/{session_id}/variables`**
- Returns: List of all variables in session namespace (kernel introspection)
- Purpose: Quick check of what exists in session
**`mcp://forge/environment/info`** (optional)
- Returns: Python version, installed packages, system info
- Purpose: Agent can verify environment capabilities
**`mcp://forge/environments/list`**
- Returns: List of available custom environments and templates
- Structure:
```json
{
"custom": [
{
"name": "my-ml-env",
"image": "mcp-forge/custom:my-ml-env",
"created": "2026-02-06T10:00:00Z",
"packages": ["numpy==1.24.0", "pandas==2.0.0"],
"size_mb": 1234
}
],
"templates": [
{
"name": "ml-basic",
"description": "Basic ML stack",
"packages": ["numpy", "pandas", "scikit-learn"]
},
{
"name": "data-science",
"description": "Data science stack",
"packages": ["numpy", "pandas", "matplotlib", "seaborn"]
}
]
}
```
**`mcp://forge/audit/operations`** (optional)
- Returns: Audit log of container operations
- Purpose: Security monitoring, debugging
#### Tools
**`execute_code`**
Execute Python or R code in an isolated container with MCP tools available.
Parameters:
```json
{
"code": "string (required) - Source code to execute",
"language": "string (optional) - 'python' (default) or 'r'",
"mcp_tools": "array (optional) - List of MCP tool names to inject",
"session_id": "string (optional) - Session ID for stateful execution (null = stateless)",
"backend": "string (optional) - 'simple' (default) or 'jupyter'",
"kernel": "string (optional) - for jupyter backend: 'python' (ipykernel) or 'r' (rkernel)",
"timeout": "integer (optional) - Max execution time in seconds (default: 300)",
"volumes": "object (optional) - Volume mount configuration",
"custom_image": "string (optional) - Custom environment name to use",
"environment": "string (optional) - Template environment name (e.g., 'ml-basic', 'data-science')"
}
```
Returns:
```json
{
"success": true,
"stdout": "string - Standard output",
"stderr": "string - Standard error",
"result": "any - Return value of last expression",
"available_tools": ["list of tools that were injected"],
"execution_time": 1.23,
"session_id": "abc123 (if stateful)"
}
```
Example:
```python
# Agent calls:
execute_python(
code="""
repos = github_search_repos(query="MCP servers", max_results=100)
high_quality = [r for r in repos['items'] if r['stars'] > 50]
print(f"Found {len(high_quality)} high-quality repos")
high_quality[:10] # Return top 10
""",
mcp_tools=["github_search_repos"],
session_id=None # stateless
)
```
**`document_state`**
Document important variables in a stateful session for later retrieval.
Parameters:
```json
{
"session_id": "string (required) - Session to document",
"variables": "object (required) - Variable name -> description mapping",
"note": "string (optional) - General note about session state",
"clear": "boolean (optional) - Clear existing documentation (default: false)"
}
```
Returns:
```json
{
"success": true,
"documented_count": 3,
"session_id": "abc123"
}
```
Example:
```python
document_state(
session_id="abc123",
variables={
"df": "Customer purchase data, 1000 rows, preprocessed and cleaned",
"model": "Trained RandomForest classifier, 87% accuracy on test set"
},
note="Ready for prediction phase"
)
```
**`build_custom_environment`**
Build a custom container image with specified Python packages. This is the **only** way to install additional packages - `pip install` is **not** allowed during code execution for security reasons.
Parameters:
```json
{
"name": "string (required) - Name for the custom environment (alphanumeric + hyphens)",
"base_image": "string (optional) - Base image to build from (default: 'python:3.11')",
"packages": "array (required) - List of Python package specifications",
"python_version": "string (optional) - Python version: '3.11', '3.12' (default: '3.11')",
"description": "string (optional) - Description of this environment"
}
```
Returns:
```json
{
"success": true,
"image_name": "mcp-forge/custom:my-ml-env",
"image_id": "sha256:abc123...",
"build_time": 45.2,
"installed_packages": ["numpy==1.24.0", "pandas==2.0.0", "scikit-learn==1.3.0"],
"cache_hit": false
}
```
Example:
```python
# Agent builds custom environment for ML work
build_custom_environment(
name="my-ml-env",
packages=[
"numpy>=1.24.0",
"pandas>=2.0.0",
"scikit-learn>=1.3.0",
"matplotlib>=3.7.0",
"seaborn>=0.12.0"
],
description="Machine learning environment with common libraries"
)
# Later, use the custom environment
execute_python(
code="import pandas as pd; df = pd.read_csv('data.csv'); ...",
custom_image="mcp-forge/custom:my-ml-env",
session_id="ml-analysis"
)
```
### 2. Execution Backends
#### Simple Backend (Default)
**Purpose:** Stateless, single-shot code execution
**Implementation:**
- Spawn Podman container with Python image
- Inject code and MCP tool functions into namespace
- Execute via `python -c`
- Capture stdout, stderr, result
- Destroy container
**Characteristics:**
- Fast startup
- No state between calls
- Ideal for data transformation tasks
- Lower memory footprint
**Use cases:**
- Data processing and filtering
- One-off computations
- Format conversions
- Quick analyses
#### Jupyter Backend
**Purpose:** Stateful, multi-step workflows
**Implementation:**
- Spawn Podman container with IPython kernel
- Keep kernel running for session lifetime
- Execute code cells via Jupyter protocol (ZMQ)
- Maintain namespace between executions
- Support rich output formats
**Characteristics:**
- State persists between calls
- Variable persistence
- Interactive workflow support
- Higher resource usage
**Use cases:**
- Multi-step data analysis
- Iterative model training
- Building up complex state
- Exploratory workflows
**Session Management:**
- Sessions identified by unique ID
- Automatic timeout after inactivity (configurable, default: 1 hour)
- Manual cleanup via session deletion
- Resource limits per session
### 3. Custom Environment Building
#### Design Philosophy
**Security Principle:** `pip install` is **NOT** allowed during code execution. All package installations must happen during a controlled build process.
**Why:**
- Prevents malicious package installation during execution
- Ensures reproducible environments
- Allows security scanning of dependencies
- Enables caching for performance
- Provides audit trail of what's installed
#### Build System: UV-Based
**Technology Choice:** [UV](https://github.com/astral-sh/uv) - ultra-fast Python package installer
**Benefits:**
- 10-100x faster than pip
- Built-in caching
- Deterministic dependency resolution
- Lockfile support for reproducibility
- Compatible with pip package specifications
#### Implementation
**Build Process:**
```python
def build_custom_environment(name, packages, base_image="python:3.11"):
"""
Build a custom container image with specified packages using UV.
Steps:
1. Generate Containerfile with UV installation
2. Create requirements specification
3. Build image with layer caching
4. Validate and tag image
5. Store metadata for reuse
"""
# Generate Containerfile
containerfile = f"""
FROM {base_image}
# Install UV (cached layer)
RUN curl -LsSf https://astral.sh/uv/install.sh | sh
ENV PATH="/root/.cargo/bin:$PATH"
# Non-root user
RUN useradd -m -u 1000 forge
USER forge
# Copy requirements (cache-friendly)
COPY requirements.txt /tmp/requirements.txt
# Install packages with UV (uses cache)
RUN uv pip install --system -r /tmp/requirements.txt
# MCP bridge client
COPY mcp_tools.py /usr/local/lib/python3.11/site-packages/
WORKDIR /workspace
"""
# Build with Podman
build_context = create_build_context(containerfile, packages)
image_id = podman.images.build(
path=build_context,
tag=f"mcp-forge/custom:{name}",
cache_from=["mcp-forge/cache:uv-base"], # UV cache layer
buildargs={
"BUILDKIT_INLINE_CACHE": "1"
}
)
return image_id
```
#### Package Cache Strategy
**Multi-Layer Caching:**
```
Layer 1: Base Python image (rarely changes)
└─ python:3.11-slim
Layer 2: UV installation (stable)
└─ UV binary + dependencies
Layer 3: Common packages (pre-cached)
└─ numpy, pandas, requests, etc.
Layer 4: User packages (specific to build)
└─ Custom requirements
```
**UV Cache Directory:**
- Mounted from host: `/var/cache/mcp-forge/uv`
- Persisted between builds
- Dramatically speeds up builds with similar dependencies
**Cache Invalidation:**
- Requirements hash changes → rebuild user layer only
- Base image updated → rebuild from base
- UV version updated → rebuild from UV layer
#### Pre-built Environment Templates
**Common Templates:**
```yaml
templates:
ml-basic:
description: "Basic ML stack"
packages:
- "numpy>=1.24.0"
- "pandas>=2.0.0"
- "scikit-learn>=1.3.0"
ml-advanced:
description: "Advanced ML with deep learning"
packages:
- "numpy>=1.24.0"
- "pandas>=2.0.0"
- "scikit-learn>=1.3.0"
- "torch>=2.0.0"
- "transformers>=4.30.0"
data-science:
description: "Data science stack"
packages:
- "numpy>=1.24.0"
- "pandas>=2.0.0"
- "matplotlib>=3.7.0"
- "seaborn>=0.12.0"
- "plotly>=5.14.0"
- "jupyter>=1.0.0"
web-scraping:
description: "Web scraping tools"
packages:
- "requests>=2.31.0"
- "beautifulsoup4>=4.12.0"
- "lxml>=4.9.0"
- "selenium>=4.10.0"
```
**Usage:**
```python
# Use a template
build_custom_environment(
name="my-ml-env",
template="ml-basic",
additional_packages=["xgboost>=2.0.0"]
)
# Or build from scratch
build_custom_environment(
name="custom-env",
packages=["specific-package==1.0.0"]
)
```
#### Build Validation
**Post-Build Checks:**
```python
def validate_built_environment(image_id):
"""
Validate that the built environment is safe and functional.
Checks:
1. All requested packages are installed
2. No malicious packages (against allowlist/blocklist)
3. Image size is within limits
4. Security scan passes (trivy/grype)
5. Python imports work
"""
# Import test
test_code = """
import sys
import json
installed = [(pkg.key, pkg.version) for pkg in __import__('pkg_resources').working_set]
print(json.dumps(installed))
"""
result = run_in_container(image_id, test_code)
installed_packages = json.loads(result.stdout)
# Security scan
scan_result = security_scan_image(image_id)
if scan_result.critical_vulns > 0:
raise SecurityError(f"Image has {scan_result.critical_vulns} critical vulnerabilities")
return {
"installed_packages": installed_packages,
"security_scan": scan_result,
"valid": True
}
```
#### Package Allowlist/Blocklist
**Security Control:**
```yaml
packages:
# Automatically allowed (common, vetted)
allowlist:
- numpy
- pandas
- scikit-learn
- matplotlib
- seaborn
- requests
- beautifulsoup4
# ... many more
# Explicitly forbidden
blocklist:
- "os-sys" # Known malicious
- "cryptography-backdoor"
# Packages with known vulnerabilities
# Requires manual approval
manual_approval:
- "*crypto*" # Cryptographic packages
- "*network*" # Network access packages
- "*subprocess*" # Process spawning
```
**Validation:**
```python
def validate_package_list(packages):
"""Check packages against allowlist/blocklist before building."""
for pkg in packages:
pkg_name = pkg.split("==")[0].split(">=")[0].split("<=")[0]
if pkg_name in BLOCKLIST:
raise SecurityError(f"Package {pkg_name} is blocked")
if not in_allowlist(pkg_name) and requires_approval(pkg_name):
raise ApprovalRequiredError(f"Package {pkg_name} requires manual approval")
```
#### Build Resource Limits
**Prevent Build Abuse:**
```python
BUILD_LIMITS = {
"max_packages": 50, # Max packages per build
"max_build_time": 600, # 10 minutes
"max_image_size": "2GB", # Final image size
"max_concurrent_builds": 3, # Per user
"build_rate_limit": {
"requests": 10,
"period": 3600 # 10 builds per hour
}
}
```
#### Environment Lifecycle
**Management:**
```python
# List user's custom environments
list_custom_environments() → [
{
"name": "my-ml-env",
"image": "mcp-forge/custom:my-ml-env",
"created": "2026-02-06T10:00:00Z",
"size_mb": 1234,
"packages": ["numpy==1.24.0", "pandas==2.0.0"],
"last_used": "2026-02-06T11:30:00Z"
}
]
# Delete unused environment
delete_custom_environment(name="old-env")
# Rebuild environment (e.g., after base image update)
rebuild_custom_environment(name="my-ml-env")
```
**Auto-cleanup:**
- Environments unused for 30 days → archived
- Archived for 90 days → deleted
- User notification before deletion
#### execute_python Integration
**Using Custom Environments:**
```python
# Updated execute_python parameters
{
"code": "...",
"custom_image": "string (optional) - Custom environment name or image tag",
"environment": "string (optional) - Template name (alternative to custom_image)",
...
}
# Examples:
execute_python(
code="import numpy as np; ...",
custom_image="my-ml-env" # Use user's custom environment
)
execute_python(
code="import pandas as pd; ...",
environment="data-science" # Use pre-built template
)
```
### 4. Container Runtime (Podman)
#### Why Podman?
- **Rootless by design:** No daemon running as root
- **Docker-compatible API:** Easy integration
- **Better security defaults:** No privileged operations needed
- **Daemonless:** Containers are child processes
#### Deployment Modes
**Development:**
```bash
# Run mcp-forge directly on host
mcp-forge serve --podman-socket /run/user/1000/podman/podman.sock
```
**Production:**
```yaml
# Run mcp-forge in container with socket mount
services:
mcp-forge:
image: mcp-forge:latest
volumes:
- /run/user/1000/podman/podman.sock:/run/podman/podman.sock:ro
environment:
- PODMAN_SOCKET=/run/podman/podman.sock
```
#### Security: Allowlist-Based Container Operations
**Allowed Operations:**
```python
ALLOWED_OPERATIONS = {
"container.create": {
"allowed_images": [
"mcp-forge/python:3.11",
"mcp-forge/python:3.12",
"mcp-forge/jupyter:latest"
],
"forbidden_params": [
"privileged",
"cap_add",
"devices",
"pid_mode",
"ipc_mode"
],
"required_params": {
"network_mode": "none", # Or restricted network
"read_only": True, # Filesystem read-only except volumes
"memory_limit": "512m",
"cpu_quota": 50000, # 50% of one CPU
"security_opt": ["no-new-privileges"],
"user": "1000:1000" # Non-root user
}
},
"container.start": {
"session_containers_only": True # Only containers we created
},
"container.stop": {
"session_containers_only": True
},
"container.remove": {
"session_containers_only": True
},
"container.logs": {
"session_containers_only": True
}
}
```
**Forbidden Operations:**
```python
FORBIDDEN_OPERATIONS = [
"container.exec", # No direct shell access
# Note: image.build IS allowed, but ONLY through build_custom_environment tool
# with validation, security scanning, and package allowlists
"image.pull", # No arbitrary image pulling (only pre-approved images)
"volume.create", # Only pre-configured volumes
"network.create", # No custom networks
"system.prune", # No system-level operations
]
# Additionally forbidden within execution containers:
EXECUTION_RESTRICTIONS = [
"subprocess.run(['pip', 'install', ...])", # No pip install during execution
"subprocess.run(['apt', 'install', ...])", # No system package installs
"import os; os.system('...')", # Restricted system calls
]
```
#### Volume Management
**Allowed Volume Patterns:**
```python
ALLOWED_VOLUME_PATTERNS = [
"/mcp-forge/sessions/{session_id}/*", # Per-session data
"/mcp-forge/shared/readonly/*", # Shared read-only data
"/mcp-forge/uploads/{session_id}/*" # User uploads
]
FORBIDDEN_MOUNT_PATHS = [
"/",
"/etc",
"/var/run/docker.sock",
"/var/run/podman/podman.sock",
"/sys",
"/proc"
]
```
#### Resource Limits (Enforced)
```python
RESOURCE_LIMITS = {
"memory": {
"default": "512m",
"max": "2g"
},
"cpu_quota": {
"default": 50000, # 50% of one CPU
"max": 100000 # 100% of one CPU
},
"timeout": {
"default": 300, # 5 minutes
"max": 1800 # 30 minutes
},
"pids_limit": 100,
"storage": {
"default": "1g",
"max": "10g"
}
}
```
### 5. MCP Tool Injection
#### Tool Discovery Flow
1. **Agent reads resource:** `mcp://forge/tools/available`
2. **MCP-Forge returns:** `["github_search_repos", "filesystem_read"]`
3. **Agent knows signatures:** Already configured in agent's tool config
4. **Agent generates code:** Using these tool names as Python functions
#### Tool Injection Implementation
**At container startup:**
```python
# MCP-Forge generates injection code
def inject_mcp_tools(container, mcp_client, tool_names):
"""
Inject MCP tools as Python functions into container namespace.
Args:
container: Podman container instance
mcp_client: Connected MCP client with available tools
tool_names: List of tool names to inject
"""
injection_code = """
# MCP Tool Functions (auto-injected)
import json
from typing import Any
def _mcp_call(tool_name: str, **kwargs) -> Any:
'''Internal: Call MCP tool via bridge'''
# This communicates with MCP-Forge server which forwards to MCP client
import socket
s = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
s.connect('/tmp/mcp-bridge.sock')
s.send(json.dumps({'tool': tool_name, 'params': kwargs}).encode())
result = json.loads(s.recv(1024*1024).decode())
s.close()
return result
"""
# Generate wrapper function for each tool
for tool_name in tool_names:
tool_schema = mcp_client.get_tool_schema(tool_name)
# Generate function signature from schema
params = tool_schema.get('inputSchema', {}).get('properties', {})
param_str = ', '.join(f"{k}: Any = None" for k in params.keys())
injection_code += f"""
def {tool_name}({param_str}):
'''
{tool_schema.get('description', '')}
Auto-generated wrapper for MCP tool: {tool_name}
'''
kwargs = {{{', '.join(f"'{k}': {k}" for k in params.keys())}}}
kwargs = {{k: v for k, v in kwargs.items() if v is not None}}
return _mcp_call('{tool_name}', **kwargs)
"""
# Write to container's Python site-packages or PYTHONSTARTUP
container.write_file('/usr/local/lib/python3.11/site-packages/mcp_tools.py',
injection_code)
# Start MCP bridge server (forwards calls to actual MCP client)
start_mcp_bridge(container, mcp_client)
```
**Alternative: Environment Variable Approach**
```python
# Set PYTHONSTARTUP to auto-import tools
container.env['PYTHONSTARTUP'] = '/mcp-forge/startup.py'
```
#### MCP Bridge Server
A lightweight server running in the MCP-Forge process that:
- Listens on Unix socket mounted into container
- Receives tool calls from container code
- Forwards to actual MCP client
- Returns results to container
This keeps the MCP client logic outside the untrusted container.
### 6. Session State Management
#### State Lifecycle
**Stateless Execution (session_id=None):**
```
Request → Create Container → Execute → Return Results → Destroy Container
```
**Stateful Execution (session_id="abc123"):**
```
First Request:
Request → Create Container → Start Kernel → Execute → Return Results
Container stays alive
Subsequent Requests:
Request → Reuse Container → Execute in Same Kernel → Return Results
Cleanup:
Timeout or Manual → Stop Container → Destroy Container
```
#### State Documentation
**Agent Workflow:**
```python
# Step 1: Load and process data
execute_python(
code="df = load_large_dataset(); df_clean = preprocess(df)",
session_id="analysis-123",
backend="jupyter"
)
# Step 2: Document important state
document_state(
session_id="analysis-123",
variables={
"df": "Raw dataset, 100k rows",
"df_clean": "Cleaned dataset, 95k rows, ready for modeling"
},
note="Data loading complete"
)
# ... conversation continues, context window rotates ...
# Step 3: Agent checks state after context rotation
state = read_resource("mcp://forge/sessions/analysis-123/state")
# Agent sees: df_clean exists and is ready
# Step 4: Continue work
execute_python(
code="model = train_model(df_clean); results = evaluate(model)",
session_id="analysis-123"
)
```
#### State Introspection
**Automatic tracking:**
- Kernel namespace inspection (all variables)
- Variable types and sizes
- Last execution timestamp
**Agent-documented:**
- Semantic descriptions
- Workflow notes
- Important variables highlighted
**Combined in resource:**
```json
{
"all_variables": ["df", "df_clean", "model", "temp", "i", "results"],
"documented_variables": {
"df_clean": "Cleaned dataset, 95k rows, ready for modeling",
"model": "Trained model, 87% accuracy",
"results": "Evaluation metrics"
},
"introspection": {
"df_clean": {"type": "DataFrame", "shape": [95000, 12], "memory_mb": 87},
"model": {"type": "RandomForestClassifier", "memory_mb": 234}
},
"note": "Ready for final predictions",
"last_updated": "2026-02-06T10:30:00Z"
}
```
## Implementation Considerations
### Technology Stack
**MCP Server:**
- Language: Python 3.11+
- Framework: `mcp` SDK (official Python implementation)
- Async: `asyncio` for concurrent operations
**Container Runtime:**
- Podman via `podman-py` library
- Fallback: Direct Podman CLI calls
**Jupyter Backend:**
- `jupyter_client` for kernel management
- ZMQ for communication
- `ipykernel` in container images
**Simple Backend:**
- Direct Python execution via `subprocess`
- Or: `python:3.11-slim` base image with `exec` entrypoint
### Pre-built Container Images
**Base Image: `mcp-forge/python:3.11`**
```dockerfile
FROM python:3.11-slim
# Non-root user
RUN useradd -m -u 1000 forge
USER forge
# Install common packages
RUN pip install --user numpy pandas requests
# MCP bridge client
COPY mcp_tools.py /usr/local/lib/python3.11/site-packages/
# Startup script
COPY startup.py /mcp-forge/startup.py
WORKDIR /workspace
```
**Jupyter Image: `mcp-forge/jupyter:latest`**
```dockerfile
FROM mcp-forge/python:3.11
USER root
RUN pip install ipykernel jupyter_client
USER forge
# IPython config for security
COPY ipython_config.py /home/forge/.ipython/profile_default/
CMD ["python", "-m", "ipykernel_launcher", "-f", "/tmp/kernel.json"]
```
### Error Handling
**Container Failures:**
- Timeout: Kill container, return timeout error
- OOM: Return memory limit error with suggestion to increase
- Crash: Return stderr and exit code
**MCP Tool Failures:**
- Tool not available: Clear error message with available tools
- Tool call error: Return tool error to agent, don't crash execution
- Network error: Retry logic for transient failures
**Security Violations:**
- Forbidden operation: Reject immediately, log to audit
- Resource limit exceeded: Terminate execution, clear error message
- Invalid volume mount: Reject with explanation
### Performance Considerations
**Container Reuse:**
- Pool of warm containers for simple backend (optional)
- Lazy cleanup of idle Jupyter sessions
**Parallel Execution:**
- Support multiple concurrent executions
- Per-session locking for stateful operations
- Configurable max concurrent containers
**Image Caching:**
- Pre-pull images on startup
- Periodic image updates (configurable schedule)
**Volume Performance:**
- Use tmpfs for ephemeral data
- Persistent volumes for session data
- Cleanup strategy for old sessions
## Security Model
### Defense in Depth
1. **Podman rootless:** No root daemon
2. **Allowlist enforcement:** Only permitted operations
3. **Read-only filesystem:** Except specific volumes
4. **Network isolation:** No internet by default
5. **Resource limits:** CPU, memory, storage, PIDs
6. **No privileged mode:** Ever
7. **Capability dropping:** Minimal capabilities
8. **User namespaces:** Non-root user in container
9. **Audit logging:** All operations logged
10. **Session isolation:** Each session has isolated volumes
### Threat Model
**Threats Mitigated:**
- Container escape → Rootless + restricted operations
- Resource exhaustion → Hard limits enforced
- Data exfiltration → Network isolation
- Privilege escalation → No privileged mode, capability restrictions
- Host filesystem access → Allowlist volumes only
**Out of Scope:**
- Side-channel attacks (Spectre, Meltdown)
- Physical security
- Supply chain attacks on base images (use trusted registries)
### Audit and Monitoring
**Logged Events:**
- Container create/start/stop/remove
- Execution requests (code hash, not full code for privacy)
- Resource limit violations
- Security policy violations
- Session lifecycle events
**Log Format:**
```json
{
"timestamp": "2026-02-06T10:30:00Z",
"event": "container.create",
"session_id": "abc123",
"image": "mcp-forge/python:3.11",
"resources": {"memory": "512m", "cpu_quota": 50000},
"success": true
}
```
## Configuration
### Server Configuration
**Example `mcp-forge.yaml`:**
```yaml
server:
host: localhost
port: 3000
podman_socket: /run/user/1000/podman/podman.sock
execution:
default_backend: simple
default_timeout: 300
max_timeout: 1800
default_memory: 512m
max_memory: 2g
default_cpu_quota: 50000
max_cpu_quota: 100000
images:
python_3_11: mcp-forge/python:3.11
python_3_12: mcp-forge/python:3.12
jupyter: mcp-forge/jupyter:latest
auto_pull: true
pull_interval: 86400 # 24 hours
sessions:
idle_timeout: 3600 # 1 hour
max_concurrent: 10
cleanup_interval: 300 # 5 minutes
volumes:
base_path: /var/lib/mcp-forge
session_quota: 1g
max_session_quota: 10g
security:
audit_log: /var/log/mcp-forge/audit.log
enforce_resource_limits: true
allow_network: false
environment_builder:
enabled: true
uv_cache_path: /var/cache/mcp-forge/uv
max_packages_per_build: 50
max_build_time: 600 # 10 minutes
max_image_size: 2147483648 # 2GB in bytes
max_concurrent_builds: 3
build_rate_limit:
requests: 10
period: 3600 # 10 builds per hour per user
# Package security
package_validation:
use_allowlist: true
allowlist_path: /etc/mcp-forge/package-allowlist.txt
blocklist_path: /etc/mcp-forge/package-blocklist.txt
require_approval_patterns:
- "*crypto*"
- "*network*"
- "*subprocess*"
# Environment lifecycle
auto_cleanup:
enabled: true
archive_after_days: 30
delete_after_days: 90
# Pre-built templates
templates:
ml-basic:
description: "Basic ML stack"
packages:
- "numpy>=1.24.0"
- "pandas>=2.0.0"
- "scikit-learn>=1.3.0"
data-science:
description: "Data science stack"
packages:
- "numpy>=1.24.0"
- "pandas>=2.0.0"
- "matplotlib>=3.7.0"
- "seaborn>=0.12.0"
- "plotly>=5.14.0"
web-scraping:
description: "Web scraping tools"
packages:
- "requests>=2.31.0"
- "beautifulsoup4>=4.12.0"
- "lxml>=4.9.0"
mcp_tools:
# MCP servers to connect to and expose
github:
command: "npx"
args: ["-y", "@modelcontextprotocol/server-github"]
env:
GITHUB_TOKEN: ${GITHUB_TOKEN}
filesystem:
command: "npx"
args: ["-y", "@modelcontextprotocol/server-filesystem", "/allowed/path"]
```
### Environment Variables
```bash
# Podman socket location
MCP_FORGE_PODMAN_SOCKET=/run/user/1000/podman/podman.sock
# Base path for volumes
MCP_FORGE_VOLUMES_PATH=/var/lib/mcp-forge
# UV cache for fast package installation
MCP_FORGE_UV_CACHE_PATH=/var/cache/mcp-forge/uv
# Security
MCP_FORGE_AUDIT_LOG=/var/log/mcp-forge/audit.log
# MCP tool credentials
GITHUB_TOKEN=ghp_xxxxx
```
## API Examples
### Example 1: Simple Data Processing
```python
# Agent discovers available tools
tools = read_resource("mcp://forge/tools/available")
# Returns: ["github_search_repos", "filesystem_read"]
# Agent generates and executes code
result = execute_python(
code="""
repos = github_search_repos(query="machine learning", max_results=100)
python_repos = [r for r in repos['items'] if r['language'] == 'Python']
top_10 = sorted(python_repos, key=lambda x: x['stars'], reverse=True)[:10]
# Return summary
[{
'name': r['name'],
'stars': r['stars'],
'url': r['html_url']
} for r in top_10]
""",
mcp_tools=["github_search_repos"],
session_id=None
)
print(result['result'])
# Returns top 10 Python ML repos
```
### Example 2: Building and Using Custom Environment
```python
# Agent needs ML libraries for analysis
# First, check available environments
envs = read_resource("mcp://forge/environments/list")
# Build custom environment if needed
build_result = build_custom_environment(
name="my-ml-analysis",
packages=[
"numpy>=1.24.0",
"pandas>=2.0.0",
"scikit-learn>=1.3.0",
"matplotlib>=3.7.0",
"xgboost>=2.0.0"
],
description="Custom ML environment for analysis project"
)
# Wait for build to complete (typically 30-60 seconds with UV cache)
print(f"Built {build_result['image_name']} in {build_result['build_time']}s")
print(f"Installed: {build_result['installed_packages']}")
# Now use the custom environment for analysis
result = execute_python(
code="""
import pandas as pd
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
import xgboost as xgb
# Load data (from volume or MCP tool)
data = filesystem_read('/data/customer_churn.csv')
df = pd.read_csv(data)
# Preprocess
X = df.drop('churn', axis=1)
y = df['churn']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)
# Train models
rf_model = RandomForestClassifier(n_estimators=100)
rf_model.fit(X_train, y_train)
rf_score = rf_model.score(X_test, y_test)
xgb_model = xgb.XGBClassifier()
xgb_model.fit(X_train, y_train)
xgb_score = xgb_model.score(X_test, y_test)
# Return results
{
'random_forest_accuracy': rf_score,
'xgboost_accuracy': xgb_score,
'best_model': 'xgboost' if xgb_score > rf_score else 'random_forest'
}
""",
custom_image="my-ml-analysis",
mcp_tools=["filesystem_read"],
session_id="churn-analysis"
)
print(result['result'])
# Returns: {'random_forest_accuracy': 0.87, 'xgboost_accuracy': 0.91, 'best_model': 'xgboost'}
```
### Example 3: Stateful Analysis with Custom Environment
```python
# Create session with custom environment
session_id = "deep-analysis"
# Step 1: Load and explore data
execute_python(
code="""
import pandas as pd
import matplotlib.pyplot as plt
df = filesystem_read_csv('/data/sales_2024.csv')
print(f"Loaded {len(df)} rows")
print(df.describe())
# Store for later
df.to_pickle('/tmp/sales_df.pkl')
""",
custom_image="data-science", # Use pre-built template
mcp_tools=["filesystem_read_csv"],
session_id=session_id,
backend="jupyter"
)
# Document what we have
document_state(
session_id=session_id,
variables={
"df": "Sales data for 2024, 50k rows, pickled to /tmp/sales_df.pkl"
},
note="Data loaded and ready for analysis"
)
# ... conversation continues, context rotates ...
# Step 2: Later, continue analysis
state = read_resource(f"mcp://forge/sessions/{session_id}/state")
# Agent sees df is available
execute_python(
code="""
import pandas as pd
df = pd.read_pickle('/tmp/sales_df.pkl')
# Analyze trends
monthly_sales = df.groupby('month')['revenue'].sum()
growth_rate = monthly_sales.pct_change().mean() * 100
{
'total_revenue': float(df['revenue'].sum()),
'avg_monthly_growth': float(growth_rate),
'top_products': df.groupby('product')['revenue'].sum().nlargest(5).to_dict()
}
""",
session_id=session_id
)
```
### Example 4: Using Pre-built Templates
```python
# Quick start with pre-built template
result = execute_python(
code="""
import requests
from bs4 import BeautifulSoup
# Scrape some data
response = requests.get('https://example.com/data')
soup = BeautifulSoup(response.content, 'lxml')
# Extract and process
data = [item.text for item in soup.find_all('div', class_='data-item')]
data[:10] # Return first 10 items
""",
environment="web-scraping", # Use template instead of building custom
session_id=None
)
print(result['result'])
```
## Implementation Roadmap
### Phase 1: Core Execution (MVP)
- Basic execute_python tool with simple backend
- Podman integration with security restrictions
- MCP tool injection
- Pre-built Python 3.11 image
### Phase 2: Stateful Execution
- Jupyter backend implementation
- Session management
- State documentation tool
- Session resources
### Phase 3: Custom Environments
- build_custom_environment tool
- UV-based package installation
- Build caching
- Package allowlist/blocklist
- Security validation
### Phase 4: Advanced Features
- Pre-built template library
- Environment lifecycle management
- Build rate limiting
- Advanced audit logging
- Multi-user support with quotas
### Phase 5: Production Hardening
- Performance optimization
- Advanced security scanning
- Comprehensive monitoring
- High availability setup
- Documentation and examples
## Multi-language Execution (Python + R)
MCP-Forge supports both Python and R execution modes:
- **Simple backend**
- `language=python` runs python shell execution (existing path)
- `language=r` runs R shell execution (new path)
- Stateless by default
- **Jupyter backend**
- `kernel=python` uses IPython kernel (existing path)
- `kernel=r` uses IRkernel (new path)
- Stateful sessions keyed by `session_id`
### Tool Injection for R
The MCP bridge protocol remains Unix-socket JSON based and language-agnostic.
R injection wrappers mirror Python wrappers at a behavioral level:
- Generate R wrapper functions per allowed MCP tool
- Forward tool calls over mounted bridge socket
- Decode JSON responses to R lists/vectors
This keeps execution containers air-gapped while preserving access to configured MCP tools.