Does the Question Really Matter? Training-Free Data Selection for Vision-Language SFT
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Peng, Shen, Huawen, Ban, Yi, Fu, Tianfan, Wang, Yanbo, Li, Yuqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reasoning BO: Enhancing Bayesian Optimization with Long-Context Reasoning Power of LLMs
von: Yang, Zhuo, et al.
Veröffentlicht: (2025)
von: Yang, Zhuo, et al.
Veröffentlicht: (2025)
KG4RecEval: Does Knowledge Graph Really Matter for Recommender Systems?
von: Zhang, Haonan, et al.
Veröffentlicht: (2024)
von: Zhang, Haonan, et al.
Veröffentlicht: (2024)
DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding
von: Liu, Yichao, et al.
Veröffentlicht: (2026)
von: Liu, Yichao, et al.
Veröffentlicht: (2026)
QCBench: Evaluating Large Language Models on Domain-Specific Quantitative Chemistry
von: Xie, Jiaqing, et al.
Veröffentlicht: (2025)
von: Xie, Jiaqing, et al.
Veröffentlicht: (2025)
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment
von: Li, Jiaxiang, et al.
Veröffentlicht: (2024)
von: Li, Jiaxiang, et al.
Veröffentlicht: (2024)
Object Attribute Matters in Visual Question Answering
von: Li, Peize, et al.
Veröffentlicht: (2023)
von: Li, Peize, et al.
Veröffentlicht: (2023)
SkillsInjector: Dynamic Skill Context Construction for LLM Agents
von: Li, Yanchao, et al.
Veröffentlicht: (2026)
von: Li, Yanchao, et al.
Veröffentlicht: (2026)
MolAct: An Agentic RL Framework for Molecular Editing and Property Optimization
von: Yang, Zhuo, et al.
Veröffentlicht: (2025)
von: Yang, Zhuo, et al.
Veröffentlicht: (2025)
ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
von: Ren, Qihan, et al.
Veröffentlicht: (2026)
von: Ren, Qihan, et al.
Veröffentlicht: (2026)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
Does More Inference-Time Compute Really Help Robustness?
von: Wu, Tong, et al.
Veröffentlicht: (2025)
von: Wu, Tong, et al.
Veröffentlicht: (2025)
Structure-based Drug Design Benchmark: Do 3D Methods Really Dominate?
von: Zheng, Kangyu, et al.
Veröffentlicht: (2024)
von: Zheng, Kangyu, et al.
Veröffentlicht: (2024)
Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?
von: Lu, Chengda, et al.
Veröffentlicht: (2025)
von: Lu, Chengda, et al.
Veröffentlicht: (2025)
ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection
von: Liu, Tao, et al.
Veröffentlicht: (2026)
von: Liu, Tao, et al.
Veröffentlicht: (2026)
SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization
von: Su, Xiaole, et al.
Veröffentlicht: (2026)
von: Su, Xiaole, et al.
Veröffentlicht: (2026)
Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations
von: Gong, Nanxu, et al.
Veröffentlicht: (2026)
von: Gong, Nanxu, et al.
Veröffentlicht: (2026)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
Object-Centric Vision Token Pruning for Vision Language Models
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
Training-Free Unsupervised Prompt for Vision-Language Models
von: Long, Sifan, et al.
Veröffentlicht: (2024)
von: Long, Sifan, et al.
Veröffentlicht: (2024)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
QCS-ADME: Quantum Circuit Search for Drug Property Prediction with Imbalanced Data and Regression Adaptation
von: Zheng, Kangyu, et al.
Veröffentlicht: (2025)
von: Zheng, Kangyu, et al.
Veröffentlicht: (2025)
Efficiency for Free: Ideal Data Are Transportable Representations
von: Sun, Peng, et al.
Veröffentlicht: (2024)
von: Sun, Peng, et al.
Veröffentlicht: (2024)
From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
von: Li, Youpeng, et al.
Veröffentlicht: (2026)
von: Li, Youpeng, et al.
Veröffentlicht: (2026)
Does Feasibility Matter? Understanding the Impact of Feasibility on Synthetic Training Data
von: Liu, Yiwen, et al.
Veröffentlicht: (2025)
von: Liu, Yiwen, et al.
Veröffentlicht: (2025)
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
von: Wei, Qianshan, et al.
Veröffentlicht: (2026)
von: Wei, Qianshan, et al.
Veröffentlicht: (2026)
MMAPG: A Training-Free Framework for Multimodal Multi-hop Question Answering via Adaptive Planning Graphs
von: Hu, Yiheng, et al.
Veröffentlicht: (2025)
von: Hu, Yiheng, et al.
Veröffentlicht: (2025)
Debunk the Myth of SFT Generalization
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2025)
DeepProtein: Deep Learning Library and Benchmark for Protein Sequence Learning
von: Xie, Jiaqing, et al.
Veröffentlicht: (2024)
von: Xie, Jiaqing, et al.
Veröffentlicht: (2024)
Are Synthetic Time-series Data Really not as Good as Real Data?
von: Fu, Fanzhe, et al.
Veröffentlicht: (2024)
von: Fu, Fanzhe, et al.
Veröffentlicht: (2024)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
von: He, Qianxi, et al.
Veröffentlicht: (2025)
von: He, Qianxi, et al.
Veröffentlicht: (2025)
Does the Skeleton-Recall Loss Really Work?
von: Arora, Devansh, et al.
Veröffentlicht: (2025)
von: Arora, Devansh, et al.
Veröffentlicht: (2025)
Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
von: Qin, Yulu, et al.
Veröffentlicht: (2025)
von: Qin, Yulu, et al.
Veröffentlicht: (2025)
OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields
von: Liu, Wanhao, et al.
Veröffentlicht: (2026)
von: Liu, Wanhao, et al.
Veröffentlicht: (2026)
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
von: Koh, Woosung, et al.
Veröffentlicht: (2026)
von: Koh, Woosung, et al.
Veröffentlicht: (2026)
RadDiff: Retrieval-Augmented Denoising Diffusion for Protein Inverse Folding
von: Han, Jin, et al.
Veröffentlicht: (2025)
von: Han, Jin, et al.
Veröffentlicht: (2025)
TOFA: Training-Free One-Shot Federated Adaptation for Vision-Language Models
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Position: How can Graphs Help Large Language Models?
von: Wang, Xiyuan, et al.
Veröffentlicht: (2026)
von: Wang, Xiyuan, et al.
Veröffentlicht: (2026)
ChemATP: A Training-Free Chemical Reasoning Framework for Large Language Models
von: Zhang, Mingxu, et al.
Veröffentlicht: (2025)
von: Zhang, Mingxu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reasoning BO: Enhancing Bayesian Optimization with Long-Context Reasoning Power of LLMs
von: Yang, Zhuo, et al.
Veröffentlicht: (2025) -
KG4RecEval: Does Knowledge Graph Really Matter for Recommender Systems?
von: Zhang, Haonan, et al.
Veröffentlicht: (2024) -
DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding
von: Liu, Yichao, et al.
Veröffentlicht: (2026) -
QCBench: Evaluating Large Language Models on Domain-Specific Quantitative Chemistry
von: Xie, Jiaqing, et al.
Veröffentlicht: (2025) -
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment
von: Li, Jiaxiang, et al.
Veröffentlicht: (2024)