Checkpoints of OPSA on different base models
Yi Ding
Tuwhy
AI & ML interests
None yet
Recent Activity
authored a paper about 8 hours ago
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement updated a model about 13 hours ago
Tuwhy/Qwen3-4B-OPSA updated a model about 13 hours ago
Tuwhy/Qwen3.5-9B-OPSAOrganizations
Sherlock
Series model of paper "Sherlock: Self-Correcting Reasoning in Vision-Language Models"
-
Tuwhy/Llama-3.2V-11B-Sherlock-SFT
Image-Text-to-Text • 11B • Updated • 8 -
Tuwhy/Llama-3.2V-11B-Sherlock-Offline
Image-Text-to-Text • 11B • Updated • 6 -
Tuwhy/Llama-3.2V-11B-Sherlock-iter1
Image-Text-to-Text • 11B • Updated • 6 -
Tuwhy/Llama-3.2V-11B-Sherlock-iter2
Image-Text-to-Text • 11B • Updated • 7 • 2
MIRage
Official model collection of paper: Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
On-Policy Self-Adaptation
Checkpoints of OPSA on different base models
Octopus
RL checkpoints of Octopus-8B and baselines of paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation
Sherlock
Series model of paper "Sherlock: Self-Correcting Reasoning in Vision-Language Models"
-
Tuwhy/Llama-3.2V-11B-Sherlock-SFT
Image-Text-to-Text • 11B • Updated • 8 -
Tuwhy/Llama-3.2V-11B-Sherlock-Offline
Image-Text-to-Text • 11B • Updated • 6 -
Tuwhy/Llama-3.2V-11B-Sherlock-iter1
Image-Text-to-Text • 11B • Updated • 6 -
Tuwhy/Llama-3.2V-11B-Sherlock-iter2
Image-Text-to-Text • 11B • Updated • 7 • 2
MIS
Official dataset collection of paper: Rethinking Bottleneck in Safety Fine-Tuning of Vision Language Models
MIRage
Official model collection of paper: Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models