Deployment Guide
Complete guide for deploying the Pneumonia Detection API to production.
Table of Contents
Local Development
Setup
# Create virtual environment
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
# Add model files
cp /path/to/svm_model.pkl models/
cp /path/to/best_densenet121.pth models/
# Run development server
python app.py
Testing
# Run test suite
python tests/test_api.py
# Manual testing
curl -X POST "http://localhost:8000/predict" \
-F "file=@test_image.jpg"
Docker Deployment
Build Docker Image
# Build image
docker build -t pneumonia-detection-api .
# Run container
docker run -p 8000:8000 pneumonia-detection-api
Using Docker Compose
# Start services
docker-compose up -d
# View logs
docker-compose logs -f
# Stop services
docker-compose down
Environment Variables
Create .env file:
# API Configuration
API_HOST=0.0.0.0
API_PORT=8000
API_WORKERS=4
# Model Configuration
SVM_MODEL_PATH=models/svm_model.pkl
CNN_MODEL_PATH=models/best_densenet121.pth
# CORS Configuration
ALLOWED_ORIGINS=https://your-frontend-domain.com
Cloud Deployment
AWS (EC2 / ECS)
EC2 Deployment:
# SSH into EC2 instance
ssh -i key.pem ubuntu@your-ec2-ip
# Install Docker
sudo apt-get update
sudo apt-get install docker.io docker-compose
# Clone repository
git clone your-repo-url
cd pneumonia-detection-backend
# Add model files
scp -i key.pem svm_model.pkl ubuntu@your-ec2-ip:~/models/
scp -i key.pem best_densenet121.pth ubuntu@your-ec2-ip:~/models/
# Run with Docker Compose
docker-compose up -d
# Configure nginx reverse proxy (optional)
sudo apt-get install nginx
# Configure /etc/nginx/sites-available/api
ECS Deployment:
# task-definition.json
{
"family": "pneumonia-detection-api",
"containerDefinitions": [
{
"name": "api",
"image": "your-ecr-repo/pneumonia-api:latest",
"portMappings": [
{
"containerPort": 8000,
"protocol": "tcp"
}
],
"memory": 2048,
"cpu": 1024
}
]
}
Google Cloud Platform (Cloud Run)
# Build and push to GCR
gcloud builds submit --tag gcr.io/PROJECT_ID/pneumonia-api
# Deploy to Cloud Run
gcloud run deploy pneumonia-api \
--image gcr.io/PROJECT_ID/pneumonia-api \
--platform managed \
--region us-central1 \
--memory 2Gi \
--cpu 2 \
--allow-unauthenticated
Heroku
# Login to Heroku
heroku login
# Create app
heroku create pneumonia-detection-api
# Add buildpacks
heroku buildpacks:set heroku/python
# Deploy
git push heroku main
# Scale dynos
heroku ps:scale web=1:standard-2x
Azure (App Service)
# Login to Azure
az login
# Create resource group
az group create --name pneumonia-api-rg --location eastus
# Create App Service plan
az appservice plan create \
--name pneumonia-api-plan \
--resource-group pneumonia-api-rg \
--sku B2 \
--is-linux
# Create web app
az webapp create \
--resource-group pneumonia-api-rg \
--plan pneumonia-api-plan \
--name pneumonia-detection-api \
--deployment-container-image-name your-dockerhub/pneumonia-api:latest
Production Considerations
1. Security
CORS Configuration:
Update app.py:
app.add_middleware(
CORSMiddleware,
allow_origins=["https://your-frontend-domain.com"], # Specific domain
allow_credentials=True,
allow_methods=["POST", "GET"],
allow_headers=["*"],
)
API Key Authentication (Optional):
from fastapi import Security, HTTPException
from fastapi.security.api_key import APIKeyHeader
API_KEY = "your-secret-api-key"
api_key_header = APIKeyHeader(name="X-API-Key")
def verify_api_key(api_key: str = Security(api_key_header)):
if api_key != API_KEY:
raise HTTPException(status_code=403, detail="Invalid API Key")
return api_key
@app.post("/predict")
async def predict(file: UploadFile, api_key: str = Depends(verify_api_key)):
# ... prediction logic
HTTPS/SSL:
- Use reverse proxy (nginx, Caddy) for SSL termination
- Use cloud provider's SSL certificates
- Use Let's Encrypt for free SSL
2. Performance Optimization
GPU Support:
Update requirements.txt:
torch==2.1.0+cu118 # CUDA 11.8 support
torchvision==0.16.0+cu118
Caching:
from functools import lru_cache
@lru_cache(maxsize=100)
def extract_features(image_hash):
# Cache frequently used features
pass
Load Balancing:
Use multiple workers:
uvicorn app:app --workers 4 --host 0.0.0.0 --port 8000
3. Rate Limiting
from slowapi import Limiter
from slowapi.util import get_remote_address
limiter = Limiter(key_func=get_remote_address)
app.state.limiter = limiter
@app.post("/predict")
@limiter.limit("10/minute")
async def predict(request: Request, file: UploadFile):
# ... prediction logic
4. File Size Limits
MAX_FILE_SIZE = 10 * 1024 * 1024 # 10 MB
@app.post("/predict")
async def predict(file: UploadFile):
contents = await file.read()
if len(contents) > MAX_FILE_SIZE:
raise HTTPException(status_code=413, detail="File too large")
# ... continue
Monitoring & Logging
1. Application Logging
Update app.py:
import logging
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s',
handlers=[
logging.FileHandler('api.log'),
logging.StreamHandler()
]
)
logger = logging.getLogger(__name__)
@app.post("/predict")
async def predict(file: UploadFile):
logger.info(f"Received prediction request for {file.filename}")
# ... prediction logic
logger.info(f"Prediction completed in {total_time}ms")
2. Prometheus Metrics
from prometheus_fastapi_instrumentator import Instrumentator
Instrumentator().instrument(app).expose(app)
3. Health Monitoring
import psutil
@app.get("/metrics")
async def metrics():
return {
"cpu_percent": psutil.cpu_percent(),
"memory_percent": psutil.virtual_memory().percent,
"disk_percent": psutil.disk_usage('/').percent
}
4. Error Tracking (Sentry)
import sentry_sdk
from sentry_sdk.integrations.fastapi import FastApiIntegration
sentry_sdk.init(
dsn="your-sentry-dsn",
integrations=[FastApiIntegration()],
)
Scaling Strategies
Horizontal Scaling
- Deploy multiple instances behind a load balancer
- Use container orchestration (Kubernetes, ECS)
- Implement session-less architecture
Vertical Scaling
- Increase CPU/RAM allocation
- Use GPU instances for faster CNN inference
- Optimize model loading (lazy loading, model quantization)
Model Optimization
- Quantization: Reduce model size and inference time
- ONNX Runtime: Convert PyTorch to ONNX for faster inference
- TensorRT: Optimize for NVIDIA GPUs
Troubleshooting
High Memory Usage
Problem: API consuming too much RAM
Solutions:
- Limit worker processes
- Implement model lazy loading
- Use model quantization
- Clear cache periodically
Slow Predictions
Problem: Predictions taking too long
Solutions:
- Use GPU if available
- Reduce image resolution (if acceptable)
- Implement request queuing
- Cache frequent predictions
Model Loading Errors
Problem: Models fail to load
Solutions:
- Verify model file paths
- Check file permissions
- Ensure sufficient disk space
- Validate model compatibility
Backup & Recovery
Model Versioning
models/
βββ v1.0/
β βββ svm_model.pkl
β βββ best_densenet121.pth
βββ v1.1/
β βββ svm_model.pkl
β βββ best_densenet121.pth
βββ current -> v1.1/
Database Backup (Optional)
If storing predictions:
# Backup PostgreSQL
pg_dump -U postgres -d predictions > backup.sql
# Restore
psql -U postgres -d predictions < backup.sql
Cost Optimization
Cloud Cost Reduction
- Use spot instances for non-critical workloads
- Auto-scaling based on traffic
- Serverless for low-traffic applications
- Reserved instances for steady workloads
Model Optimization
- Use smaller models for lower costs
- Implement request batching
- Cache predictions when possible
Compliance & Privacy
HIPAA Compliance (if applicable)
- Use encrypted storage for images
- Implement audit logging
- Use BAA-compliant cloud providers
- Anonymize patient data
GDPR Compliance
- Implement data retention policies
- Allow users to delete their data
- Provide data export functionality
- Obtain explicit consent
Recommended Architecture
Internet
β
Load Balancer (nginx/AWS ALB)
β
API Instances (Docker containers)
βββ Model Cache (Redis)
βββ Prediction Queue (RabbitMQ/SQS)
βββ Logging (CloudWatch/ELK)
β
Storage
βββ Model Files (S3/GCS)
βββ Database (PostgreSQL/MongoDB)
Support
For deployment assistance:
- Check documentation: README.md
- Review logs:
docker-compose logs - Test endpoints:
python tests/test_api.py
Version: 1.0.0
Last Updated: January 2026