# Track A Solution ## Competition Context Track A is the wireless network troubleshooting and optimization track of the Telco Troubleshooting Agent challenge. The task is to solve 5G drive-test troubleshooting scenarios by selecting the best optimization action from a list of candidate choices. For each scenario, the solution must produce one final prediction and record the generated reasoning/completion trace. In the final evaluation phase, the organizer provides the private test scenarios and deploys the sandbox tool server locally. The submitted code is expected to run without manual intervention and write the required output files under `result/`. The expected final output shape is: ```text result/ traces.json results.csv runtime.json ``` `results.csv` must contain: ```csv scenario_id,prediction ``` ## Packaged Files This submission directory is designed to be zipped as `code.zip`. ```text README.md SOLUTION.md run.py main.py train.py requirements.txt data/ Phase_1/ train.json src/ __init__.py model_core.py models/ deploy.sh model_v4_bundle.pkl result/ results.csv traces.json runtime.json ``` The `result/` directory contains the currently staged outputs copied from the existing Track A run. The evaluator can regenerate the same required filenames by running `run.py`. ## Solution Overview The solution combines a trained auxiliary candidate ranker with an agentic reasoning layer. The main components are: - `models/model_v4_bundle.pkl`: trained Track A auxiliary model bundle. - `data/Phase_1/train.json`: labelled Phase 1 training data used to reproduce the auxiliary model. - `train.py`: training entrypoint. It cross-validates the template classifier and candidate selector, trains final models on all labelled data, and writes the final bundle to `models/model_v4_bundle.pkl`. - `src/model_core.py`: shared feature extraction, option parsing, model scoring, server hydration, and output writing utilities. - `main.py`: original Track A inference entrypoint. It loads scenarios, optionally hydrates placeholder telemetry from the Track A tool server, loads the model bundle, produces predictions, and records traces. - `run.py`: submission wrapper. It calls `main.py`, then converts the raw Track A output into the Phase 3-required `result/results.csv`, `result/traces.json`, and `result/runtime.json`. - `models/deploy.sh`: vLLM deployment helper for the Qwen3.5-35B-A3B base model. ## Architecture The code is organized as a small, reproducible training and inference stack: ```text Training path: data/Phase_1/train.json -> train.py -> cross-validation metrics and experiment artifacts -> final template classifier and candidate selector -> models/model_v4_bundle.pkl Inference path: private test.json -> run.py -> main.py -> optional sandbox hydration -> feature extraction and option parsing -> template classifier -> candidate option selector -> bounded agentic reranker -> raw result.csv, debug.json, traces.json -> run.py output normalizer -> result/results.csv, result/traces.json, result/runtime.json ``` ### Submission Wrapper `run.py` is the competition-facing entrypoint. It keeps the package output in the required Phase 3 shape while preserving the existing Track A inference code. It: - Resolves paths relative to the submission directory. - Calls `main.py` in a subprocess. - Writes raw intermediate files under `result/_raw_track_a/`. - Converts the raw Track A `ID,Track A,Track B` CSV into `scenario_id,prediction`. - Normalizes `traces.json` so every trace has `scenario_id`, `completion_id`, and `completion`. - Creates `runtime.json` from the per-scenario `execution_time_seconds` in the traces. - Deletes only the temporary `_raw_track_a/` folder after conversion. ### Inference Controller `main.py` owns the batch inference loop. It: - Loads the test scenarios. - Detects whether scenario data contains placeholder text instructing the code to use the API. - Hydrates missing telemetry from the Track A sandbox when needed. - Loads `models/model_v4_bundle.pkl`. - Builds an OpenAI-compatible client if an API key is available. - Solves each scenario with `run_agent_with_trained_model`. - Records ordered results, debug rows, and trace rows. - Checkpoints after each scenario by default. The per-scenario `solve_one` function is also where runtime is measured with `time.perf_counter()`. That timing is written into each trace as `execution_time_seconds`, and `run.py` later converts it into `runtime.json`. ### Sandbox Data Hydration The code treats API calls as deterministic data gathering, not as free-form agent actions. If a scenario contains placeholder telemetry, `hydrate_scenario_from_server` fetches: - User-plane drive-test data. - Network configuration data. - KPI or traffic data. - MR data. - Signaling-plane event logs when required. Those calls are appended to the trace in compact form. This keeps the LLM from inventing telemetry and makes the agent operate over structured evidence. ## Model And Candidate-Ranker Design The auxiliary model bundle contains two fitted scikit-learn style pipelines: 1. A template classifier. 2. A selector model. ### Template Classifier The template classifier predicts the likely action pattern for a scenario. A template is an ordered action summary such as: ```text dec_power|tilt_down inc_power|azimuth neighbor dec_power|tilt_down|inc_a3|inc_a3 ``` The action vocabulary is derived from the candidate option text and includes: ```text server insufficient pdcch inc_power dec_power tilt_down tilt_up azimuth inc_a3 dec_a3 a2a5 neighbor ``` This gives the system a compact prior over "what kind of fix is likely needed" before choosing the exact candidate option ID. ### Selector Model The selector scores every candidate option, conditioned on a template. It is not limited to options that exactly match the top template. At prediction time the code evaluates the top five templates and uses the selector to score candidate options under those templates. The selector features combine: - Scenario-level telemetry features. - Whether the task is single-answer or multiple-answer. - The candidate option action type. - The candidate option target cell. - Parsed option amount and unit, such as dB, dBm, degrees, or SYM. - A2 versus A5 threshold subtype. - PDCCH symbol count. - Target-cell serving and neighbor evidence. - Target-cell configuration values such as power, downtilt, thresholds, and PDCCH symbols. - Geometry features such as azimuth error and desired downtilt. - Proposed-change features, such as post-change power margin or azimuth amount error. The final ML prior is selected by combining the selector rank score with a small template probability bonus: ```text candidate_score = selector_rank_score + 0.05 * template_probability ``` For single-answer tasks, the top ranked candidate is selected. For multiple-answer tasks, the selector tries to satisfy the predicted action counts while preserving diversity across action, target cell, and threshold subtype. ## Training And Reproducibility Flow The submission includes the labelled Phase 1 training set: ```text data/Phase_1/train.json ``` The training entrypoint is: ```text train.py ``` To rebuild the auxiliary model from scratch, run this command from the submission directory: ```bash python train.py \ --train_path data/Phase_1/train.json \ --out models/model_v4_bundle.pkl \ --experiment_name lgbm_v4 \ --n_jobs -1 ``` The training script performs the following steps: 1. Loads the labelled Phase 1 training scenarios. 2. Builds scenario-level features for the template classifier. 3. Builds all-option selector rows so every candidate option can be scored. 4. Cross-validates the template and selector models. 5. Writes fold metrics, overall metrics, OOF predictions, split manifests, split JSON files, logs, and metadata under `results/experiments/`. 6. Trains final models on all labelled training data. 7. Saves the final experiment bundle under the experiment directory. 8. Copies the final trained model bundle to `models/model_v4_bundle.pkl`. That final `models/model_v4_bundle.pkl` file is the model artifact loaded by `main.py` during inference. ## Agent Design Choices The agent is intentionally narrow and bounded. The design goal is to let the LLM inspect compact evidence and correct obvious candidate-choice mistakes without giving it enough freedom to destabilize the run. ### One Tool Only The only exposed agent tool is: ```text run_trained_ml_model ``` That tool returns: - The ML prior prediction. - The selected labels. - The predicted template and top templates. - The original options. - Compact evidence for each option. - Ambiguity groups where similar options differ by amount, target, subtype, or PDCCH symbol count. The LLM does not call raw sandbox endpoints. Data collection happens before the agent loop, and the agent sees curated evidence. ### Forced First Tool Call On the first step, `run_agent_with_trained_model` sets `tool_choice` to force `run_trained_ml_model`. This ensures every agent answer is grounded in the model prior and the structured evidence object. After that first tool call, the LLM may return strict JSON: ```json {"final_answer":"C13","used_tool":true,"reason":"short"} ``` ### Conservative Reranking The system prompt tells the LLM to treat the trained model as a strong prior, not as an unchangeable answer. The LLM is allowed to rerank only when the returned option evidence indicates a better: - Target cell. - Action subtype, such as A2 versus A5. - Numerical magnitude, such as azimuth degrees or threshold delta. - PDCCH symbol value. This mirrors the candidate-ranker augmentation idea: Phase 2-style problems often require choosing among options that share the same broad action but differ in target, amount, or subtype. The selector exposes those distinctions, and the agent acts as a final evidence-aware sanity check. ### Validation And Fallbacks The agent loop is guarded by: - `max_steps`. - `max_tool_calls`. - Per-question timeout. - Duplicate tool-call detection. - Strict JSON response validation. - Valid candidate ID filtering. - Cardinality checks for single-answer and multiple-answer tasks. If the LLM is unavailable, times out, returns invalid answers, calls duplicate tools, or raises an exception, the code falls back to the direct ML prior. This means every scenario can still produce a deterministic prediction. ## Solution Architecture And Data Flow The following diagram shows the end-to-end data flow implemented by `run.py`, `main.py`, and `run_agent_with_trained_model`. ```text +-----------------+ | Start: scenario | +-----------------+ | v +-----------------------------+ | Load and normalize scenario | +-----------------------------+ | v +------------------------+ | Placeholder telemetry? | +------------------------+ | yes | no v v +-----------------------------+ +--------------------------+ | Hydrate from Track A sandbox | --> | Build prediction context | +-----------------------------+ +--------------------------+ | v +-----------------------------------------------+ | Extract scenario, cell, geometry, and option | | features | +-----------------------------------------------+ | v +-----------------------------------------------+ | Template classifier: top action templates | +-----------------------------------------------+ | v +-----------------------------------------------+ | Selector model: score options for top | | templates | +-----------------------------------------------+ | v +-----------------------------------------------+ | ML prior labels and option evidence | +-----------------------------------------------+ | v +-----------------------+ | LLM client available? | +-----------------------+ | yes | no v v +--------------------------------------+ +--------------------+ | Agent: forced run_trained_ml_model | | Direct ML fallback | | tool call | +--------------------+ +--------------------------------------+ | | | v | +--------------------------------------+ | | Tool result: prior, top templates, | | | options, ambiguity groups | | +--------------------------------------+ | | | v | +--------------------------------------+ | | Agent reranks and returns strict JSON| | +--------------------------------------+ | | | v | +-------------------------+ | | Valid IDs and cardinality? | +-------------------------+ | | yes | no | v v | +--------------+ +--------------+ | | Final labels | | Budget left? | | +--------------+ +--------------+ | ^ | yes | no | | v v | | +----------------+ | | | | Reprompt for | | | | | strict JSON | | | | +----------------+ | | | | | | +---------------+-----------+-------+ | v +------------------------------------------------+ | main.py writes raw result, debug, and traces | +------------------------------------------------+ | v +------------------------------------------------+ | run.py normalizes required output files | +------------------------------------------------+ | v +------------------------------------------------+ | End: result/results.csv, traces.json, | | runtime.json | +------------------------------------------------+ ``` The state carried through this graph is: - `scenario`: original or hydrated scenario record. - `context`: extracted scenario features, cell stats, geometry stats, options, actions, and semantics. - `ml_result`: model prediction, top templates, selector evidence, and ambiguity groups. - `messages`: bounded LLM conversation state. - `tool_trace`: compact record of data hydration and trained-model tool calls. - `labels`: final normalized C-label prediction. - `debug`: template probabilities, fallback reason, timeout status, and whether the agent changed the ML prior. ## Runtime Flow 1. The organizer deploys the Qwen3.5-35B-A3B base model with vLLM, or uses `models/deploy.sh` as a deployment template. 2. The organizer starts the Track A sandbox tool server locally. 3. `run.py` is called with the private Track A test file. 4. `run.py` delegates inference to `main.py`. 5. `main.py` loads the auxiliary model bundle and solves each scenario. 6. `main.py` writes raw intermediate output under `result/_raw_track_a/` and writes raw traces to `result/traces.json`. 7. `run.py` converts the raw `result.csv` into the required `results.csv` format. 8. `run.py` derives `runtime.json` from the per-scenario execution timings stored in the traces. 9. `run.py` removes the temporary `result/_raw_track_a/` directory. The final result directory contains only the required competition output files. ## Environment Install dependencies from the submission directory: ```bash pip install -r requirements.txt ``` The solution expects: - Python 3.10 or newer. - A local OpenAI-compatible vLLM endpoint for Qwen3.5-35B-A3B. - The Track A sandbox tool server at `https://localhost:8081/no`, unless overridden. - No token authentication for the Track A sandbox endpoint. Default model endpoint: ```text http://localhost:8001/v1 ``` Default served model name: ```text Qwen3.5-35B-A3B ``` ## Model Deployment The base model weights are not included in the submission. The organizer-provided or locally mounted Qwen3.5-35B-A3B weights should be deployed with vLLM. Example: ```bash BASE_MODEL_PATH=/path/to/Qwen3.5-35B-A3B bash models/deploy.sh ``` The deployment script serves the model through an OpenAI-compatible API on port `8001` by default. ## Running The Submission From inside this submission directory: ```bash python run.py --input /path/to/test.json --output result ``` Useful explicit form: ```bash python run.py \ --input /path/to/test.json \ --output result \ --server_url https://localhost:8081/no \ --model_url http://localhost:8001/v1 \ --model_name Qwen3.5-35B-A3B ``` `run.py` sets a placeholder API key automatically if no OpenAI-compatible API key environment variable is present, which is sufficient for many local vLLM deployments. ## Output Files ### `result/results.csv` Contains one final Track A prediction per scenario: ```csv scenario_id,prediction ``` Predictions are normalized choice labels such as: ```text C13 C5|C9 ``` ### `result/traces.json` Contains the recorded completions and tool-call context for each scenario. Each trace includes at least: ```text scenario_id completion_id completion ``` Additional fields such as question text, selected prediction, tool calls, and execution time are retained for auditability. ### `result/runtime.json` Contains per-scenario runtime in seconds: ```json [ { "scenario_id": "example-scenario-id", "runtime_seconds": 12.34 } ] ``` Runtime values are derived from the scenario-level timing recorded by the inference code. ## Staged Output Verification The included staged outputs were copied from: ```text hf_dataset/Track A/results/ ``` The staged output folder has been checked for: - `results.csv` present. - `traces.json` present. - `runtime.json` present. - 500 predictions in `results.csv`. - 500 trace entries. - 500 runtime entries. - No blank Track A predictions. - Matching scenario IDs across results, traces, and runtime files. ## Notes The auxiliary model bundle is packaged as `models/model_v4_bundle.pkl`. This is the model artifact used by the current Track A pipeline. The submission wrapper does not train models, download files, or require internet access during inference when the base model, dependencies, and sandbox services are already available locally.