Models and dataset for "Parameter-Efficient Multimodal Instruction Tuning for Romanian Vision–Language Models"