--- title: Openenv Claw1 emoji: 🚀 colorFrom: indigo colorTo: gray sdk: docker app_port: 7860 pinned: false --- # Last-Mile Delivery Optimization (OpenEnv) OpenEnv-compatible benchmark environment for city-scale last-mile dispatch under realistic operating constraints. ## Motivation Urban delivery systems face a difficult planning problem: dispatchers must complete as many deliveries as possible while navigating congestion, road blockages, and battery limits. Poor routing or action sequencing creates delays, SLA violations for high-priority orders, and unnecessary energy cost. This environment models that real-world challenge in a deterministic grid world so policies can be compared fairly across difficulty levels. ## What This Environment Simulates - Multi-order pickup and delivery workflow. - Static obstacles, dynamic obstacles, and traffic penalties. - Priority-aware delivery pressure (especially in hard mode). - Battery constraints and charging behavior in hard mode. - Deterministic seeded task generation for reproducible evaluation. ## Architecture Diagram ```mermaid flowchart LR U[OpenEnv Evaluator or User] -->|HTTP| API[FastAPI API Layer\napp/main.py + app/routes.py] API -->|reset or step or state| ENV[LastMileDeliveryEnvironment\nenv/environment.py] ENV --> SIM[DeliverySimulator\nenv/simulator.py] SIM -->|observation, reward, done, info| API API -->|baseline endpoint rollout| BASE[BaselineGreedyAgent\nbaseline/baseline_agent.py] BASE --> ENV API -->|grader endpoint| GRADER[DeliveryEpisodeGrader\ngrader/grader.py] GRADER --> REPORT[GradeReport\nscore + metrics + safeguards] TASKS[tasks/easy.py + tasks/medium.py + tasks/hard.py\n+ tasks/registry.py] -->|task config| ENV TASKS -->|success_condition| GRADER TASKS -->|metadata| API ``` Flow summary: - The API orchestrates environment state transitions and stores trajectory steps. - The baseline endpoint runs a fresh rollout with the built-in greedy agent. - The grader endpoint scores the recorded trajectory against task success conditions. - Task definitions provide deterministic configuration and grading targets. ## Action Space The action payload is a strict JSON object with exactly four fields. | Field | Type | Required | Allowed Values | Notes | |---|---|---|---|---| | move | string or null | Yes | up, down, left, right, stay, null | Movement action; stay maps to no movement | | accept_order | integer or null | Yes | order index or null | When set to n, maps to internal order_n | | deliver_order | boolean | Yes | true or false | Deliver current order when true | | wait | boolean | Yes | true or false | Explicit wait action when true | Validation is strict. Exactly one action intent must be active per step. ## Observation Space Each step returns an observation object with agent, order, and map state. | Field | Type | Description | |---|---|---| | grid_width, grid_height | integer | Map dimensions | | agent_location | object {x, y} | Current agent position | | pending_orders | array of Order | Orders not yet delivered | | current_order | Order or null | Active accepted/picked-up order | | obstacles | array of Position | Static + dynamic blocked cells | | dynamic_obstacles | array of Position | Current dynamic blocked cells | | traffic_zones | array of TrafficZone | Cells with extra movement cost | | charging_stations | array of Position | Recharge cells | | battery_level | integer or null | Active in battery-enabled tasks | | step_count | integer | Elapsed step count | | total_reward | float | Cumulative episode reward | ## Tasks All tasks are deterministic and expose explicit success conditions. ### Easy - Description: simple single-order delivery on a compact map. - Configuration: 6x6 grid, 1 order, max_steps=60. - Constraints: no obstacles, no traffic, no battery, no priority mix. - Success condition: completion_rate=1.0, max_steps=60, invalid_action_rate_max=0.12. ### Medium - Description: multi-order delivery with static obstacles and traffic friction. - Configuration: 10x10 grid, 3 to 4 orders, max_steps=140. - Constraints: static obstacles + traffic enabled. - Success condition: completion_rate_min=0.85, max_steps=140, invalid_action_rate_max=0.08. ### Hard - Description: high-pressure dispatch with tight battery and SLA constraints. - Configuration: 16x16 grid, 6 to 8 orders, max_steps=165. - Constraints: static + dynamic obstacles, traffic, high-priority orders, multi-stop delivery, battery + recharge, delay pressure. - Success condition: completion_rate_min=0.90, high_priority_on_time_rate_min=0.88, battery_depletion=false, max_steps=165, invalid_action_rate_max=0.04. ## Reward Design The reward is dense and combines task completion signals with operational efficiency: - Base step penalty: -1.0 each step. - Delivery reward: +50.0 on successful delivery. - Destination milestone: +10.0 when pickup/drop destination is reached. - Invalid action penalty: -20.0. - Traffic movement penalty: -2.0 when entering traffic cells. - Progress shaping: signed distance-based reward clipped by configured bounds. - Priority delivery bonus: +15.0 (high) or +5.0 (low). - Recharge reward: +2.0 when battery is restored at charging stations. - Delay penalty: scaled negative penalty for overdue active orders. - Battery failure penalty: -30.0 when battery depletes to zero. Design goal: favor correct and efficient order handling while discouraging invalid and wasteful behavior. ## Grader Logic The grader is deterministic and aligned to each task's success_condition fields. Main metrics: - completion_rate: delivered_orders / total_orders. - high_priority_on_time_rate: on-time high-priority deliveries. - efficiency_ratio: normalized by max_steps target. - invalid_action_rate: invalid_actions / steps_taken. Scoring components: - Completion component: highest weight. - Priority SLA component: enabled when high_priority_on_time_rate_min exists. - Efficiency component: step-budget performance. - Penalty component: reduces score when invalid action rate exceeds target. Edge-case safeguards include: - Zero-step and zero-order episode handling. - Invalid success-condition value fallback and clamping. - Division-safe invalid-action and efficiency calculations. - Optional battery depletion penalty when the task requires no depletion. ## Setup ### Local 1. Create and activate a virtual environment. 2. Install dependencies. 3. Run the API server on port 7860. ```bash python -m venv .venv source .venv/bin/activate pip install -r requirements.txt uvicorn app.main:app --host 0.0.0.0 --port 7860 ``` ### Docker ```bash docker build -t delivery-openenv:latest . docker run --rm -p 7860:7860 \ -e HF_TOKEN="example-provider-key" \ -e MODEL_NAME="gpt-4o-mini" \ -e API_BASE_URL="https://api.openai.com/v1" \ -e OPENAI_USE_MODEL="1" \ delivery-openenv:latest ``` ## API Usage Base URL: ```text http://0.0.0.0:7860 ``` ### Health ```bash curl -s http://0.0.0.0:7860/health ``` ### Frontend Results Dashboard Open the interactive dashboard in your browser: ```bash open http://0.0.0.0:7860/ui ``` The dashboard lets you run `reset`, `step`, `state`, `grader`, and `baseline` calls and inspect live payloads and metrics. ### List Tasks and Action Schema ```bash curl -s http://0.0.0.0:7860/tasks ``` ### Reset Environment ```bash curl -s -X POST http://0.0.0.0:7860/reset \ -H "Content-Type: application/json" \ -d '{"task":"easy","seed":101}' ``` ### Reset From Real-World Scenario Use this endpoint when you want to inject externally curated map and order state instead of task-generated layouts. ```bash curl -s -X POST http://0.0.0.0:7860/reset_from_scenario \ -H "Content-Type: application/json" \ -d '{ "seed": 99, "success_condition": { "completion_rate_min": 1.0, "max_steps": 20, "invalid_action_rate_max": 0.20, "battery_depletion": false }, "scenario": { "width": 6, "height": 6, "max_steps": 20, "agent_start": {"x": 0, "y": 0}, "orders": [ { "order_id": "r1", "pickup": {"x": 1, "y": 0}, "dropoff": {"x": 2, "y": 0}, "delivery_locations": [{"x": 2, "y": 0}], "priority": "high", "created_step": 0 } ], "obstacles": [{"x": 4, "y": 4}], "dynamic_obstacles": [{"x": 4, "y": 5}], "traffic_zones": [{"location": {"x": 3, "y": 0}, "extra_cost": 3}], "charging_stations": [{"x": 0, "y": 0}], "battery_profile": { "enabled": true, "capacity": 10, "recharge_rate": 4, "initial_level": 7 } } }' ``` Contract summary for `scenario`: - `width`, `height`, `max_steps`: grid and episode budget. - `agent_start`: starting position. - `orders`: list of orders with pickup/dropoff/delivery locations and priority. - `obstacles`, `dynamic_obstacles`: blocked cells. - `traffic_zones`: movement-cost cells (`extra_cost >= 1`). - `charging_stations`: recharge cells. - `battery_profile`: battery enablement, capacity, recharge rate, and optional initial level. - `success_condition` (top-level): optional custom grader thresholds for this scenario. ### Step ```bash curl -s -X POST http://0.0.0.0:7860/step \ -H "Content-Type: application/json" \ -d '{"move":null,"accept_order":null,"deliver_order":false,"wait":true}' ``` ### State ```bash curl -s http://0.0.0.0:7860/state ``` ### Grade Current Episode ```bash curl -s http://0.0.0.0:7860/grader ``` ### Baseline Rollout Evaluation ```bash curl -s "http://0.0.0.0:7860/baseline?task=easy" ``` ## Baseline Scores (Reproducible Inference Run) Measured with `inference.py --all-tasks --seed 42 --json-summary` using deterministic fallback mode (`OPENAI_USE_MODEL=0`). | Task | Steps | Success | Score | |---|---:|---:|---:| | easy | 6 | true | 0.9700 | | medium | 140 | false | 0.0001 | | hard | 55 | false | 0.0667 | Average score (seed=42): **0.3456** ## Real-World Validation We tested the environment under realistic scenarios: - Priority-based delivery selection. - Traffic-aware routing decisions. - Battery-constrained navigation. Results indicate that evaluated agents can exhibit behavior aligned with real-world logistics systems under these constraints. ## OpenEnv Manifest Project metadata and runtime wiring are defined in openenv.yaml, including: - runtime entrypoints, - environment class, - task registry loader, - grader class, - action and observation schema declarations. ## Baseline Inference Script The baseline runner uses the OpenAI Python client and supports the required submission variables: - `HF_TOKEN`: API key/token used as OpenAI client `api_key`. - `MODEL_NAME`: model identifier used for inference requests. - `API_BASE_URL`: optional OpenAI-compatible endpoint. Backward-compatible aliases are also supported: `OPENAI_API_KEY`, `OPENAI_MODEL`, and `OPENAI_BASE_URL`. ```bash export HF_TOKEN="your-api-key" export MODEL_NAME="gpt-4o-mini" export API_BASE_URL="https://api.openai.com/v1" export OPENAI_USE_MODEL="1" # Reproducible baseline score across easy/medium/hard tasks. python inference.py --all-tasks --seed 42 --json-summary ``` Notes: - `OPENAI_USE_MODEL=1` enables remote model calls. - `OPENAI_USE_MODEL=0` keeps deterministic heuristic fallback while still validating OpenAI client wiring. - `--seed` controls reproducibility for baseline comparisons. ## Trainable Agent (Q-Learning) This repository now includes a trainable tabular Q-learning agent for interactive learning on the OpenEnv loop. Training script: ```bash python scripts/train_q_agent.py --task easy --episodes 2000 --eval-episodes 200 --seed 42 --output models/q_agent_easy.json ``` What it does: - Trains through repeated `reset()` / `step()` episodes. - Learns action values for a compact state representation. - Evaluates the learned policy and prints success/completion metrics. - Saves a reusable model artifact (`models/q_agent_easy.json`). Learner runtime components: - Agent implementation: `baseline/trained_q_agent.py` - Trainer entrypoint: `scripts/train_q_agent.py` ## Validation Commands ```bash ./scripts/validate.sh .venv/bin/openenv validate python -m pytest -q ``` ## Test Command ```bash python -m pytest -q ``` ## Hugging Face Spaces (Docker) 1. Create a Docker Space. 2. Push this repository. 3. Set variables/secrets: HF_TOKEN, MODEL_NAME, API_BASE_URL, OPENAI_USE_MODEL. 4. Ensure exposed port is 7860. 5. Rebuild and verify /health. Additional deployment notes are available in HF_SPACES_DEPLOYMENT.md.