Vision and Language, Video, Multimodal Learning
Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation