Spaces:
Paused
Paused
| # FoundationPose Quick Start Guide | |
| ## Overview | |
| This Hugging Face Space provides two modes: | |
| 1. **Placeholder Mode** (default) - Returns empty results, useful for testing the API without GPU requirements | |
| 2. **Real Mode** - Uses actual FoundationPose model for 6D pose estimation (requires GPU and model weights) | |
| ## Testing Locally (Placeholder Mode) | |
| The easiest way to test the API structure: | |
| ```bash | |
| cd foundationpose | |
| pip install -r requirements.txt | |
| python app.py | |
| ``` | |
| Visit http://localhost:7860 to see the UI. | |
| ## Deploying to Hugging Face Spaces | |
| ### Option 1: Placeholder Mode (No Setup Required) | |
| Just push to your Space: | |
| ```bash | |
| cd foundationpose | |
| git add . | |
| git commit -m "Deploy FoundationPose Space" | |
| git push | |
| ``` | |
| The Space will run in placeholder mode by default. This is useful for: | |
| - Testing the API structure | |
| - Developing client integrations | |
| - Demos without GPU costs | |
| ### Option 2: Real FoundationPose (Requires Setup) | |
| **Step 1: Clone FoundationPose Repository** | |
| ```bash | |
| # Inside your local foundationpose directory | |
| git clone https://github.com/NVlabs/FoundationPose.git | |
| ``` | |
| **Step 2: Download Model Weights** | |
| Download weights from the official Google Drive: | |
| https://drive.google.com/drive/folders/1GCyGE-LbFGgRC-FuGsF3a1zeBuzsQ1Da | |
| Extract to: | |
| ``` | |
| foundationpose/weights/ | |
| ├── 2023-10-28-18-33-37/ (refiner weights) | |
| └── 2024-01-11-20-02-45/ (scorer weights) | |
| ``` | |
| **Step 3: Add Weights to Git LFS** | |
| ```bash | |
| git lfs install | |
| git lfs track "weights/**/*.pth" | |
| git lfs track "weights/**/*.ckpt" | |
| git add .gitattributes | |
| git add weights/ | |
| git commit -m "Add model weights" | |
| ``` | |
| **Step 4: Enable Real Mode** | |
| Add to your Space settings (or use .env file locally): | |
| ``` | |
| USE_REAL_MODEL=true | |
| ``` | |
| **Step 5: Push to HF** | |
| ```bash | |
| git push | |
| ``` | |
| ## Using the API | |
| ### Python Client | |
| ```python | |
| from foundationpose.client import FoundationPoseClient | |
| import cv2 | |
| import numpy as np | |
| # Initialize client | |
| client = FoundationPoseClient("https://gpue-foundationpose.hf.space") | |
| # Load reference images | |
| ref_images = [] | |
| for i in range(1, 16): | |
| img = cv2.imread(f"reference/image_{i:03d}.jpg") | |
| img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) | |
| ref_images.append(img) | |
| # Register object | |
| client.initialize("target_cube", ref_images) | |
| # Estimate pose | |
| query_img = cv2.imread("query.jpg") | |
| query_img = cv2.cvtColor(query_img, cv2.COLOR_BGR2RGB) | |
| poses = client.estimate_pose("target_cube", query_img) | |
| print(f"Detected {len(poses)} poses") | |
| for pose in poses: | |
| print(f"Position: {pose['position']}") | |
| print(f"Orientation: {pose['orientation']}") | |
| print(f"Confidence: {pose['confidence']}") | |
| ``` | |
| ### Direct HTTP API | |
| ```bash | |
| # Initialize | |
| curl -X POST https://gpue-foundationpose.hf.space/api/initialize \ | |
| -H "Content-Type: application/json" \ | |
| -d '{ | |
| "object_id": "target_cube", | |
| "reference_images_b64": ["'$(base64 -w 0 ref1.jpg)'", "'$(base64 -w 0 ref2.jpg)'"], | |
| "camera_intrinsics": "{\"fx\": 500, \"fy\": 500, \"cx\": 320, \"cy\": 240}" | |
| }' | |
| # Estimate | |
| curl -X POST https://gpue-foundationpose.hf.space/api/estimate \ | |
| -H "Content-Type: application/json" \ | |
| -d '{ | |
| "object_id": "target_cube", | |
| "query_image_b64": "'$(base64 -w 0 query.jpg)'" | |
| }' | |
| ``` | |
| ## Integration with robot-ml Training | |
| Update `/training/nova_sim_trainer/perception/foundation_pose_wrapper.py`: | |
| ```python | |
| from foundationpose.client import FoundationPoseClient | |
| class FoundationPoseWrapper(PoseEstimator): | |
| def __init__(self, api_url: str, tracked_objects: List[Dict], **kwargs): | |
| super().__init__() | |
| self.client = FoundationPoseClient(api_url) | |
| # Initialize each tracked object | |
| for obj_config in tracked_objects: | |
| object_id = obj_config["object_id"] | |
| ref_dir = Path(obj_config["reference_images_dir"]) | |
| # Load reference images | |
| ref_images = [] | |
| for img_path in sorted(ref_dir.glob("*.jpg")): | |
| img = cv2.imread(str(img_path)) | |
| img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) | |
| ref_images.append(img) | |
| # Register object | |
| self.client.initialize(object_id, ref_images) | |
| logger.info(f"Registered {object_id} with {len(ref_images)} images") | |
| def estimate_poses(self, frame, camera_intrinsics, scene_objects): | |
| # Call API for pose estimation | |
| poses = self.client.estimate_pose( | |
| self.tracked_objects[0]["object_id"], # For now, single object | |
| frame, | |
| camera_intrinsics | |
| ) | |
| # Convert to DetectedPose format | |
| return [DetectedPose(**pose) for pose in poses] | |
| ``` | |
| Update `observations.yaml`: | |
| ```yaml | |
| perception: | |
| enabled: true | |
| model: foundation_pose | |
| api_url: https://gpue-foundationpose.hf.space | |
| tracked_objects: | |
| - object_id: target_cube | |
| reference_images_dir: ./perception/reference/target_cube | |
| ``` | |
| ## Performance Tips | |
| ### Cold Start Latency | |
| - First request takes 15-30s (GPU allocation + model loading) | |
| - Subsequent requests: 0.5-2s | |
| ### Keeping GPU Warm | |
| Send periodic keep-alive requests: | |
| ```python | |
| import time | |
| import threading | |
| def keep_warm(): | |
| while True: | |
| try: | |
| client.estimate_pose("target_cube", dummy_image) | |
| except: | |
| pass | |
| time.sleep(60) # Every minute | |
| threading.Thread(target=keep_warm, daemon=True).start() | |
| ``` | |
| ### Batch Processing | |
| For recorded episodes, process all frames in one session: | |
| ```python | |
| # Initialize once | |
| client.initialize("target_cube", ref_images) | |
| # Process all frames | |
| poses_list = [] | |
| for frame in frames: | |
| poses = client.estimate_pose("target_cube", frame) | |
| poses_list.append(poses) | |
| ``` | |
| ## Troubleshooting | |
| ### Space shows "Placeholder mode" | |
| - Set `USE_REAL_MODEL=true` in Space secrets | |
| - Verify weights are uploaded correctly | |
| - Check Space logs for errors | |
| ### "Model weights not found" | |
| - Ensure weights are in `weights/` directory | |
| - Check git-lfs tracked files: `git lfs ls-files` | |
| - Re-upload if needed | |
| ### GPU timeout | |
| - Increase `@spaces.GPU(duration=X)` in app.py | |
| - Reduce image resolution | |
| - Process fewer reference images | |
| ### Out of memory | |
| - Use lower resolution images | |
| - Process fewer objects simultaneously | |
| - Request more GPU resources in Space settings | |
| ## Cost Optimization | |
| ZeroGPU is free but has usage limits: | |
| - **Development**: Use placeholder mode | |
| - **Testing**: Enable real mode for specific tests only | |
| - **Production**: Consider dedicated GPU deployment (RunPod, Modal, etc.) | |
| ## Next Steps | |
| 1. Test locally in placeholder mode | |
| 2. Upload weights for real mode | |
| 3. Integrate with robot-ml training pipeline | |
| 4. Monitor GPU usage and costs | |
| 5. Optimize batch processing for your use case | |
| ## Support | |
| - **Issues**: https://github.com/gpuschel/robot-ml/issues | |
| - **FoundationPose**: https://github.com/NVlabs/FoundationPose | |
| - **HF Spaces**: https://huggingface.co/docs/hub/spaces | |