openenv-claw1 / README.md
vishal harkal
Deploy full OpenEnv API instead of starter app
f104717
|
Raw History Blame Contribute Delete
12.5 kB
metadata
title: Openenv Claw1
emoji: 🚀
colorFrom: indigo
colorTo: gray
sdk: docker
app_port: 7860
pinned: false

Last-Mile Delivery Optimization (OpenEnv)

OpenEnv-compatible benchmark environment for city-scale last-mile dispatch under realistic operating constraints.

Motivation

Urban delivery systems face a difficult planning problem: dispatchers must complete as many deliveries as possible while navigating congestion, road blockages, and battery limits. Poor routing or action sequencing creates delays, SLA violations for high-priority orders, and unnecessary energy cost.

This environment models that real-world challenge in a deterministic grid world so policies can be compared fairly across difficulty levels.

What This Environment Simulates

  • Multi-order pickup and delivery workflow.
  • Static obstacles, dynamic obstacles, and traffic penalties.
  • Priority-aware delivery pressure (especially in hard mode).
  • Battery constraints and charging behavior in hard mode.
  • Deterministic seeded task generation for reproducible evaluation.

Architecture Diagram

flowchart LR
    U[OpenEnv Evaluator or User] -->|HTTP| API[FastAPI API Layer\napp/main.py + app/routes.py]

    API -->|reset or step or state| ENV[LastMileDeliveryEnvironment\nenv/environment.py]
    ENV --> SIM[DeliverySimulator\nenv/simulator.py]
    SIM -->|observation, reward, done, info| API

    API -->|baseline endpoint rollout| BASE[BaselineGreedyAgent\nbaseline/baseline_agent.py]
    BASE --> ENV

    API -->|grader endpoint| GRADER[DeliveryEpisodeGrader\ngrader/grader.py]
    GRADER --> REPORT[GradeReport\nscore + metrics + safeguards]

    TASKS[tasks/easy.py + tasks/medium.py + tasks/hard.py\n+ tasks/registry.py] -->|task config| ENV
    TASKS -->|success_condition| GRADER
    TASKS -->|metadata| API

Flow summary:

  • The API orchestrates environment state transitions and stores trajectory steps.
  • The baseline endpoint runs a fresh rollout with the built-in greedy agent.
  • The grader endpoint scores the recorded trajectory against task success conditions.
  • Task definitions provide deterministic configuration and grading targets.

Action Space

The action payload is a strict JSON object with exactly four fields.

Field Type Required Allowed Values Notes
move string or null Yes up, down, left, right, stay, null Movement action; stay maps to no movement
accept_order integer or null Yes order index or null When set to n, maps to internal order_n
deliver_order boolean Yes true or false Deliver current order when true
wait boolean Yes true or false Explicit wait action when true

Validation is strict. Exactly one action intent must be active per step.

Observation Space

Each step returns an observation object with agent, order, and map state.

Field Type Description
grid_width, grid_height integer Map dimensions
agent_location object {x, y} Current agent position
pending_orders array of Order Orders not yet delivered
current_order Order or null Active accepted/picked-up order
obstacles array of Position Static + dynamic blocked cells
dynamic_obstacles array of Position Current dynamic blocked cells
traffic_zones array of TrafficZone Cells with extra movement cost
charging_stations array of Position Recharge cells
battery_level integer or null Active in battery-enabled tasks
step_count integer Elapsed step count
total_reward float Cumulative episode reward

Tasks

All tasks are deterministic and expose explicit success conditions.

Easy

  • Description: simple single-order delivery on a compact map.
  • Configuration: 6x6 grid, 1 order, max_steps=60.
  • Constraints: no obstacles, no traffic, no battery, no priority mix.
  • Success condition: completion_rate=1.0, max_steps=60, invalid_action_rate_max=0.12.

Medium

  • Description: multi-order delivery with static obstacles and traffic friction.
  • Configuration: 10x10 grid, 3 to 4 orders, max_steps=140.
  • Constraints: static obstacles + traffic enabled.
  • Success condition: completion_rate_min=0.85, max_steps=140, invalid_action_rate_max=0.08.

Hard

  • Description: high-pressure dispatch with tight battery and SLA constraints.
  • Configuration: 16x16 grid, 6 to 8 orders, max_steps=165.
  • Constraints: static + dynamic obstacles, traffic, high-priority orders, multi-stop delivery, battery + recharge, delay pressure.
  • Success condition: completion_rate_min=0.90, high_priority_on_time_rate_min=0.88, battery_depletion=false, max_steps=165, invalid_action_rate_max=0.04.

Reward Design

The reward is dense and combines task completion signals with operational efficiency:

  • Base step penalty: -1.0 each step.
  • Delivery reward: +50.0 on successful delivery.
  • Destination milestone: +10.0 when pickup/drop destination is reached.
  • Invalid action penalty: -20.0.
  • Traffic movement penalty: -2.0 when entering traffic cells.
  • Progress shaping: signed distance-based reward clipped by configured bounds.
  • Priority delivery bonus: +15.0 (high) or +5.0 (low).
  • Recharge reward: +2.0 when battery is restored at charging stations.
  • Delay penalty: scaled negative penalty for overdue active orders.
  • Battery failure penalty: -30.0 when battery depletes to zero.

Design goal: favor correct and efficient order handling while discouraging invalid and wasteful behavior.

Grader Logic

The grader is deterministic and aligned to each task's success_condition fields.

Main metrics:

  • completion_rate: delivered_orders / total_orders.
  • high_priority_on_time_rate: on-time high-priority deliveries.
  • efficiency_ratio: normalized by max_steps target.
  • invalid_action_rate: invalid_actions / steps_taken.

Scoring components:

  • Completion component: highest weight.
  • Priority SLA component: enabled when high_priority_on_time_rate_min exists.
  • Efficiency component: step-budget performance.
  • Penalty component: reduces score when invalid action rate exceeds target.

Edge-case safeguards include:

  • Zero-step and zero-order episode handling.
  • Invalid success-condition value fallback and clamping.
  • Division-safe invalid-action and efficiency calculations.
  • Optional battery depletion penalty when the task requires no depletion.

Setup

Local

  1. Create and activate a virtual environment.
  2. Install dependencies.
  3. Run the API server on port 7860.
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --host 0.0.0.0 --port 7860

Docker

docker build -t delivery-openenv:latest .
docker run --rm -p 7860:7860 \
    -e HF_TOKEN="example-provider-key" \
    -e MODEL_NAME="gpt-4o-mini" \
    -e API_BASE_URL="https://api.openai.com/v1" \
    -e OPENAI_USE_MODEL="1" \
    delivery-openenv:latest

API Usage

Base URL:

http://0.0.0.0:7860

Health

curl -s http://0.0.0.0:7860/health

Frontend Results Dashboard

Open the interactive dashboard in your browser:

open http://0.0.0.0:7860/ui

The dashboard lets you run reset, step, state, grader, and baseline calls and inspect live payloads and metrics.

List Tasks and Action Schema

curl -s http://0.0.0.0:7860/tasks

Reset Environment

curl -s -X POST http://0.0.0.0:7860/reset \
    -H "Content-Type: application/json" \
    -d '{"task":"easy","seed":101}'

Reset From Real-World Scenario

Use this endpoint when you want to inject externally curated map and order state instead of task-generated layouts.

curl -s -X POST http://0.0.0.0:7860/reset_from_scenario \
    -H "Content-Type: application/json" \
    -d '{
      "seed": 99,
      "success_condition": {
        "completion_rate_min": 1.0,
        "max_steps": 20,
        "invalid_action_rate_max": 0.20,
        "battery_depletion": false
      },
      "scenario": {
        "width": 6,
        "height": 6,
        "max_steps": 20,
        "agent_start": {"x": 0, "y": 0},
        "orders": [
          {
            "order_id": "r1",
            "pickup": {"x": 1, "y": 0},
            "dropoff": {"x": 2, "y": 0},
            "delivery_locations": [{"x": 2, "y": 0}],
            "priority": "high",
            "created_step": 0
          }
        ],
        "obstacles": [{"x": 4, "y": 4}],
        "dynamic_obstacles": [{"x": 4, "y": 5}],
        "traffic_zones": [{"location": {"x": 3, "y": 0}, "extra_cost": 3}],
        "charging_stations": [{"x": 0, "y": 0}],
        "battery_profile": {
          "enabled": true,
          "capacity": 10,
          "recharge_rate": 4,
          "initial_level": 7
        }
      }
    }'

Contract summary for scenario:

  • width, height, max_steps: grid and episode budget.
  • agent_start: starting position.
  • orders: list of orders with pickup/dropoff/delivery locations and priority.
  • obstacles, dynamic_obstacles: blocked cells.
  • traffic_zones: movement-cost cells (extra_cost >= 1).
  • charging_stations: recharge cells.
  • battery_profile: battery enablement, capacity, recharge rate, and optional initial level.
  • success_condition (top-level): optional custom grader thresholds for this scenario.

Step

curl -s -X POST http://0.0.0.0:7860/step \
    -H "Content-Type: application/json" \
    -d '{"move":null,"accept_order":null,"deliver_order":false,"wait":true}'

State

curl -s http://0.0.0.0:7860/state

Grade Current Episode

curl -s http://0.0.0.0:7860/grader

Baseline Rollout Evaluation

curl -s "http://0.0.0.0:7860/baseline?task=easy"

Baseline Scores (Reproducible Inference Run)

Measured with inference.py --all-tasks --seed 42 --json-summary using deterministic fallback mode (OPENAI_USE_MODEL=0).

Task Steps Success Score
easy 6 true 0.9700
medium 140 false 0.0001
hard 55 false 0.0667

Average score (seed=42): 0.3456

Real-World Validation

We tested the environment under realistic scenarios:

  • Priority-based delivery selection.
  • Traffic-aware routing decisions.
  • Battery-constrained navigation.

Results indicate that evaluated agents can exhibit behavior aligned with real-world logistics systems under these constraints.

OpenEnv Manifest

Project metadata and runtime wiring are defined in openenv.yaml, including:

  • runtime entrypoints,
  • environment class,
  • task registry loader,
  • grader class,
  • action and observation schema declarations.

Baseline Inference Script

The baseline runner uses the OpenAI Python client and supports the required submission variables:

  • HF_TOKEN: API key/token used as OpenAI client api_key.
  • MODEL_NAME: model identifier used for inference requests.
  • API_BASE_URL: optional OpenAI-compatible endpoint.

Backward-compatible aliases are also supported: OPENAI_API_KEY, OPENAI_MODEL, and OPENAI_BASE_URL.

export HF_TOKEN="your-api-key"
export MODEL_NAME="gpt-4o-mini"
export API_BASE_URL="https://api.openai.com/v1"
export OPENAI_USE_MODEL="1"

# Reproducible baseline score across easy/medium/hard tasks.
python inference.py --all-tasks --seed 42 --json-summary

Notes:

  • OPENAI_USE_MODEL=1 enables remote model calls.
  • OPENAI_USE_MODEL=0 keeps deterministic heuristic fallback while still validating OpenAI client wiring.
  • --seed controls reproducibility for baseline comparisons.

Trainable Agent (Q-Learning)

This repository now includes a trainable tabular Q-learning agent for interactive learning on the OpenEnv loop.

Training script:

python scripts/train_q_agent.py --task easy --episodes 2000 --eval-episodes 200 --seed 42 --output models/q_agent_easy.json

What it does:

  • Trains through repeated reset() / step() episodes.
  • Learns action values for a compact state representation.
  • Evaluates the learned policy and prints success/completion metrics.
  • Saves a reusable model artifact (models/q_agent_easy.json).

Learner runtime components:

  • Agent implementation: baseline/trained_q_agent.py
  • Trainer entrypoint: scripts/train_q_agent.py

Validation Commands

./scripts/validate.sh
.venv/bin/openenv validate
python -m pytest -q

Test Command

python -m pytest -q

Hugging Face Spaces (Docker)

  1. Create a Docker Space.
  2. Push this repository.
  3. Set variables/secrets: HF_TOKEN, MODEL_NAME, API_BASE_URL, OPENAI_USE_MODEL.
  4. Ensure exposed port is 7860.
  5. Rebuild and verify /health.

Additional deployment notes are available in HF_SPACES_DEPLOYMENT.md.