Spaces:
Sleeping
Download README.md from vishalharkal/openenv-claw1: direct link, hf CLI and curl.
- Browser
- Download file 12.5 kB
-
https://huggingface.co/spaces/vishalharkal/openenv-claw1/resolve/main/README.md
- Command line
-
hf download hf://spaces/vishalharkal/openenv-claw1/README.md
-
curl -L -o README.md https://huggingface.co/spaces/vishalharkal/openenv-claw1/resolve/main/README.md
title: Openenv Claw1
emoji: 🚀
colorFrom: indigo
colorTo: gray
sdk: docker
app_port: 7860
pinned: false
Last-Mile Delivery Optimization (OpenEnv)
OpenEnv-compatible benchmark environment for city-scale last-mile dispatch under realistic operating constraints.
Motivation
Urban delivery systems face a difficult planning problem: dispatchers must complete as many deliveries as possible while navigating congestion, road blockages, and battery limits. Poor routing or action sequencing creates delays, SLA violations for high-priority orders, and unnecessary energy cost.
This environment models that real-world challenge in a deterministic grid world so policies can be compared fairly across difficulty levels.
What This Environment Simulates
- Multi-order pickup and delivery workflow.
- Static obstacles, dynamic obstacles, and traffic penalties.
- Priority-aware delivery pressure (especially in hard mode).
- Battery constraints and charging behavior in hard mode.
- Deterministic seeded task generation for reproducible evaluation.
Architecture Diagram
flowchart LR
U[OpenEnv Evaluator or User] -->|HTTP| API[FastAPI API Layer\napp/main.py + app/routes.py]
API -->|reset or step or state| ENV[LastMileDeliveryEnvironment\nenv/environment.py]
ENV --> SIM[DeliverySimulator\nenv/simulator.py]
SIM -->|observation, reward, done, info| API
API -->|baseline endpoint rollout| BASE[BaselineGreedyAgent\nbaseline/baseline_agent.py]
BASE --> ENV
API -->|grader endpoint| GRADER[DeliveryEpisodeGrader\ngrader/grader.py]
GRADER --> REPORT[GradeReport\nscore + metrics + safeguards]
TASKS[tasks/easy.py + tasks/medium.py + tasks/hard.py\n+ tasks/registry.py] -->|task config| ENV
TASKS -->|success_condition| GRADER
TASKS -->|metadata| API
Flow summary:
- The API orchestrates environment state transitions and stores trajectory steps.
- The baseline endpoint runs a fresh rollout with the built-in greedy agent.
- The grader endpoint scores the recorded trajectory against task success conditions.
- Task definitions provide deterministic configuration and grading targets.
Action Space
The action payload is a strict JSON object with exactly four fields.
| Field | Type | Required | Allowed Values | Notes |
|---|---|---|---|---|
| move | string or null | Yes | up, down, left, right, stay, null | Movement action; stay maps to no movement |
| accept_order | integer or null | Yes | order index or null | When set to n, maps to internal order_n |
| deliver_order | boolean | Yes | true or false | Deliver current order when true |
| wait | boolean | Yes | true or false | Explicit wait action when true |
Validation is strict. Exactly one action intent must be active per step.
Observation Space
Each step returns an observation object with agent, order, and map state.
| Field | Type | Description |
|---|---|---|
| grid_width, grid_height | integer | Map dimensions |
| agent_location | object {x, y} | Current agent position |
| pending_orders | array of Order | Orders not yet delivered |
| current_order | Order or null | Active accepted/picked-up order |
| obstacles | array of Position | Static + dynamic blocked cells |
| dynamic_obstacles | array of Position | Current dynamic blocked cells |
| traffic_zones | array of TrafficZone | Cells with extra movement cost |
| charging_stations | array of Position | Recharge cells |
| battery_level | integer or null | Active in battery-enabled tasks |
| step_count | integer | Elapsed step count |
| total_reward | float | Cumulative episode reward |
Tasks
All tasks are deterministic and expose explicit success conditions.
Easy
- Description: simple single-order delivery on a compact map.
- Configuration: 6x6 grid, 1 order, max_steps=60.
- Constraints: no obstacles, no traffic, no battery, no priority mix.
- Success condition: completion_rate=1.0, max_steps=60, invalid_action_rate_max=0.12.
Medium
- Description: multi-order delivery with static obstacles and traffic friction.
- Configuration: 10x10 grid, 3 to 4 orders, max_steps=140.
- Constraints: static obstacles + traffic enabled.
- Success condition: completion_rate_min=0.85, max_steps=140, invalid_action_rate_max=0.08.
Hard
- Description: high-pressure dispatch with tight battery and SLA constraints.
- Configuration: 16x16 grid, 6 to 8 orders, max_steps=165.
- Constraints: static + dynamic obstacles, traffic, high-priority orders, multi-stop delivery, battery + recharge, delay pressure.
- Success condition: completion_rate_min=0.90, high_priority_on_time_rate_min=0.88, battery_depletion=false, max_steps=165, invalid_action_rate_max=0.04.
Reward Design
The reward is dense and combines task completion signals with operational efficiency:
- Base step penalty: -1.0 each step.
- Delivery reward: +50.0 on successful delivery.
- Destination milestone: +10.0 when pickup/drop destination is reached.
- Invalid action penalty: -20.0.
- Traffic movement penalty: -2.0 when entering traffic cells.
- Progress shaping: signed distance-based reward clipped by configured bounds.
- Priority delivery bonus: +15.0 (high) or +5.0 (low).
- Recharge reward: +2.0 when battery is restored at charging stations.
- Delay penalty: scaled negative penalty for overdue active orders.
- Battery failure penalty: -30.0 when battery depletes to zero.
Design goal: favor correct and efficient order handling while discouraging invalid and wasteful behavior.
Grader Logic
The grader is deterministic and aligned to each task's success_condition fields.
Main metrics:
- completion_rate: delivered_orders / total_orders.
- high_priority_on_time_rate: on-time high-priority deliveries.
- efficiency_ratio: normalized by max_steps target.
- invalid_action_rate: invalid_actions / steps_taken.
Scoring components:
- Completion component: highest weight.
- Priority SLA component: enabled when high_priority_on_time_rate_min exists.
- Efficiency component: step-budget performance.
- Penalty component: reduces score when invalid action rate exceeds target.
Edge-case safeguards include:
- Zero-step and zero-order episode handling.
- Invalid success-condition value fallback and clamping.
- Division-safe invalid-action and efficiency calculations.
- Optional battery depletion penalty when the task requires no depletion.
Setup
Local
- Create and activate a virtual environment.
- Install dependencies.
- Run the API server on port 7860.
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --host 0.0.0.0 --port 7860
Docker
docker build -t delivery-openenv:latest .
docker run --rm -p 7860:7860 \
-e HF_TOKEN="example-provider-key" \
-e MODEL_NAME="gpt-4o-mini" \
-e API_BASE_URL="https://api.openai.com/v1" \
-e OPENAI_USE_MODEL="1" \
delivery-openenv:latest
API Usage
Base URL:
http://0.0.0.0:7860
Health
curl -s http://0.0.0.0:7860/health
Frontend Results Dashboard
Open the interactive dashboard in your browser:
open http://0.0.0.0:7860/ui
The dashboard lets you run reset, step, state, grader, and baseline calls and inspect live payloads and metrics.
List Tasks and Action Schema
curl -s http://0.0.0.0:7860/tasks
Reset Environment
curl -s -X POST http://0.0.0.0:7860/reset \
-H "Content-Type: application/json" \
-d '{"task":"easy","seed":101}'
Reset From Real-World Scenario
Use this endpoint when you want to inject externally curated map and order state instead of task-generated layouts.
curl -s -X POST http://0.0.0.0:7860/reset_from_scenario \
-H "Content-Type: application/json" \
-d '{
"seed": 99,
"success_condition": {
"completion_rate_min": 1.0,
"max_steps": 20,
"invalid_action_rate_max": 0.20,
"battery_depletion": false
},
"scenario": {
"width": 6,
"height": 6,
"max_steps": 20,
"agent_start": {"x": 0, "y": 0},
"orders": [
{
"order_id": "r1",
"pickup": {"x": 1, "y": 0},
"dropoff": {"x": 2, "y": 0},
"delivery_locations": [{"x": 2, "y": 0}],
"priority": "high",
"created_step": 0
}
],
"obstacles": [{"x": 4, "y": 4}],
"dynamic_obstacles": [{"x": 4, "y": 5}],
"traffic_zones": [{"location": {"x": 3, "y": 0}, "extra_cost": 3}],
"charging_stations": [{"x": 0, "y": 0}],
"battery_profile": {
"enabled": true,
"capacity": 10,
"recharge_rate": 4,
"initial_level": 7
}
}
}'
Contract summary for scenario:
width,height,max_steps: grid and episode budget.agent_start: starting position.orders: list of orders with pickup/dropoff/delivery locations and priority.obstacles,dynamic_obstacles: blocked cells.traffic_zones: movement-cost cells (extra_cost >= 1).charging_stations: recharge cells.battery_profile: battery enablement, capacity, recharge rate, and optional initial level.success_condition(top-level): optional custom grader thresholds for this scenario.
Step
curl -s -X POST http://0.0.0.0:7860/step \
-H "Content-Type: application/json" \
-d '{"move":null,"accept_order":null,"deliver_order":false,"wait":true}'
State
curl -s http://0.0.0.0:7860/state
Grade Current Episode
curl -s http://0.0.0.0:7860/grader
Baseline Rollout Evaluation
curl -s "http://0.0.0.0:7860/baseline?task=easy"
Baseline Scores (Reproducible Inference Run)
Measured with inference.py --all-tasks --seed 42 --json-summary using deterministic fallback mode (OPENAI_USE_MODEL=0).
| Task | Steps | Success | Score |
|---|---|---|---|
| easy | 6 | true | 0.9700 |
| medium | 140 | false | 0.0001 |
| hard | 55 | false | 0.0667 |
Average score (seed=42): 0.3456
Real-World Validation
We tested the environment under realistic scenarios:
- Priority-based delivery selection.
- Traffic-aware routing decisions.
- Battery-constrained navigation.
Results indicate that evaluated agents can exhibit behavior aligned with real-world logistics systems under these constraints.
OpenEnv Manifest
Project metadata and runtime wiring are defined in openenv.yaml, including:
- runtime entrypoints,
- environment class,
- task registry loader,
- grader class,
- action and observation schema declarations.
Baseline Inference Script
The baseline runner uses the OpenAI Python client and supports the required submission variables:
HF_TOKEN: API key/token used as OpenAI clientapi_key.MODEL_NAME: model identifier used for inference requests.API_BASE_URL: optional OpenAI-compatible endpoint.
Backward-compatible aliases are also supported: OPENAI_API_KEY, OPENAI_MODEL, and OPENAI_BASE_URL.
export HF_TOKEN="your-api-key"
export MODEL_NAME="gpt-4o-mini"
export API_BASE_URL="https://api.openai.com/v1"
export OPENAI_USE_MODEL="1"
# Reproducible baseline score across easy/medium/hard tasks.
python inference.py --all-tasks --seed 42 --json-summary
Notes:
OPENAI_USE_MODEL=1enables remote model calls.OPENAI_USE_MODEL=0keeps deterministic heuristic fallback while still validating OpenAI client wiring.--seedcontrols reproducibility for baseline comparisons.
Trainable Agent (Q-Learning)
This repository now includes a trainable tabular Q-learning agent for interactive learning on the OpenEnv loop.
Training script:
python scripts/train_q_agent.py --task easy --episodes 2000 --eval-episodes 200 --seed 42 --output models/q_agent_easy.json
What it does:
- Trains through repeated
reset()/step()episodes. - Learns action values for a compact state representation.
- Evaluates the learned policy and prints success/completion metrics.
- Saves a reusable model artifact (
models/q_agent_easy.json).
Learner runtime components:
- Agent implementation:
baseline/trained_q_agent.py - Trainer entrypoint:
scripts/train_q_agent.py
Validation Commands
./scripts/validate.sh
.venv/bin/openenv validate
python -m pytest -q
Test Command
python -m pytest -q
Hugging Face Spaces (Docker)
- Create a Docker Space.
- Push this repository.
- Set variables/secrets: HF_TOKEN, MODEL_NAME, API_BASE_URL, OPENAI_USE_MODEL.
- Ensure exposed port is 7860.
- Rebuild and verify /health.
Additional deployment notes are available in HF_SPACES_DEPLOYMENT.md.