Papers
arxiv:2608.28958

CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions

Published on Aug 29
ยท Submitted by
Patrick (Tsung-Han) Wu
on Sep 1
Authors:
,
,
,
,
,
,

Abstract

CoVA-SFT is a multimodal reasoning dataset that teaches language models to interleave text and visual abstractions via structured reasoning steps, improving performance on visual reasoning benchmarks.

Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkward prose. Although architectural solutions exist to process visual inputs, the community lacks a massive, multi-step, self-corrected dataset to teach models how to build and maintain internal visual workspaces when solving purely textual reasoning problems. To address this limitation, we introduce CoVA-SFT, a highly structured corpus of 51.9K samples containing over 222K multimodal reasoning steps across 5 distinct layout families and 17 complex tasks, and CoVA-Bench, a companion benchmark of 1,700 held-out test samples spanning the same tasks for reproducible evaluation. By providing explicit rationale formulations, agentic renderings, and verification loops, CoVA-SFT teaches multimodal language models to interleave text and visual abstractions. We validate the dataset by demonstrating that models fine-tuned on CoVA-SFT outperform all interleaved CoT baselines by more than 2x on average on CoVA-Bench, though they still fall short of strong text-only CoT baselines, highlighting open challenges for future work.

Community

Paper submitter

What if thinking itself could be visual, even when the problem isn't?

We release CoVA-SFT: 51.9K visual reasoning traces (222K rendered steps) for text-only problems across 17 tasks, plus CoVA-Bench.

๐Ÿ“ EMNLP 2026 Findings
๐Ÿ“„ https://arxiv.org/abs/2608.28958
๐Ÿค— https://huggingface.co/datasets/tsunghanwu/cova
fig2_X

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.28958 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.28958 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.