Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
Fuente:
arXiv
Guardado en:
| Autores principales: | Gkountouras, John, Titov, Ivan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025)
por: Yim, Wen-wai, et al.
Publicado: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
por: Ma, Yueen, et al.
Publicado: (2024)
por: Ma, Yueen, et al.
Publicado: (2024)
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)
por: Cai, Weibin, et al.
Publicado: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
por: Chen, Kewei, et al.
Publicado: (2026)
por: Chen, Kewei, et al.
Publicado: (2026)
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
por: Malikussaid, et al.
Publicado: (2026)
por: Malikussaid, et al.
Publicado: (2026)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
por: Radosky, Lukas, et al.
Publicado: (2026)
por: Radosky, Lukas, et al.
Publicado: (2026)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
por: Chen, Kewei, et al.
Publicado: (2025)
por: Chen, Kewei, et al.
Publicado: (2025)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
por: Tu, Songjun, et al.
Publicado: (2025)
por: Tu, Songjun, et al.
Publicado: (2025)
Does CLIP perceive art the same way we do?
por: Asperti, Andrea, et al.
Publicado: (2025)
por: Asperti, Andrea, et al.
Publicado: (2025)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
por: Zhang, Junwen, et al.
Publicado: (2025)
por: Zhang, Junwen, et al.
Publicado: (2025)
RACAS: Controlling Diverse Robots With a Single Agentic System
por: Ashley, Dylan R., et al.
Publicado: (2026)
por: Ashley, Dylan R., et al.
Publicado: (2026)
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
por: Nwokocha, Caleb Princewill
Publicado: (2022)
por: Nwokocha, Caleb Princewill
Publicado: (2022)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
por: Cui, Hejie, et al.
Publicado: (2024)
por: Cui, Hejie, et al.
Publicado: (2024)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
por: Badshah, Sher, et al.
Publicado: (2025)
por: Badshah, Sher, et al.
Publicado: (2025)
How much do LLMs learn from negative examples?
por: Hamdan, Shadi, et al.
Publicado: (2025)
por: Hamdan, Shadi, et al.
Publicado: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
por: Estevanell-Valladares, Ernesto L., et al.
Publicado: (2025)
por: Estevanell-Valladares, Ernesto L., et al.
Publicado: (2025)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
por: Fu, Tianyu, et al.
Publicado: (2024)
por: Fu, Tianyu, et al.
Publicado: (2024)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
por: Sáez, Arnau Igualde, et al.
Publicado: (2025)
por: Sáez, Arnau Igualde, et al.
Publicado: (2025)
Linguistic Collapse: Neural Collapse in (Large) Language Models
por: Wu, Robert, et al.
Publicado: (2024)
por: Wu, Robert, et al.
Publicado: (2024)
Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study
por: Tikhonov, Alexey, et al.
Publicado: (2025)
por: Tikhonov, Alexey, et al.
Publicado: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
por: Farzulla, Murad
Publicado: (2026)
por: Farzulla, Murad
Publicado: (2026)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
por: Kaiser, Daniel, et al.
Publicado: (2025)
por: Kaiser, Daniel, et al.
Publicado: (2025)
Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
por: Ouyang, Rongxin, et al.
Publicado: (2024)
por: Ouyang, Rongxin, et al.
Publicado: (2024)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
por: Viveiros, André G., et al.
Publicado: (2025)
por: Viveiros, André G., et al.
Publicado: (2025)
Do Reasoning Models Enhance Embedding Models?
por: Chan, Wun Yu, et al.
Publicado: (2026)
por: Chan, Wun Yu, et al.
Publicado: (2026)
Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
por: Kuz, Mykola, et al.
Publicado: (2025)
por: Kuz, Mykola, et al.
Publicado: (2025)
Cost-Effective Attention Mechanisms for Low Resource Settings: Necessity & Sufficiency of Linear Transformations
por: Hosseini, Peyman, et al.
Publicado: (2024)
por: Hosseini, Peyman, et al.
Publicado: (2024)
TextTeacher: What Can Language Teach About Images?
por: Nauen, Tobias Christian, et al.
Publicado: (2026)
por: Nauen, Tobias Christian, et al.
Publicado: (2026)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
por: Huo, Dongjie, et al.
Publicado: (2026)
por: Huo, Dongjie, et al.
Publicado: (2026)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
por: Papyan, Narek, et al.
Publicado: (2024)
por: Papyan, Narek, et al.
Publicado: (2024)
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
por: Sela, Omer
Publicado: (2026)
por: Sela, Omer
Publicado: (2026)
Harnessing non-adversarial robustness in large language models
por: Zhou, Qinghua, et al.
Publicado: (2026)
por: Zhou, Qinghua, et al.
Publicado: (2026)
LLM-supported document separation for printed reviews from zbMATH Open
por: Pluzhnikov, Ivan, et al.
Publicado: (2026)
por: Pluzhnikov, Ivan, et al.
Publicado: (2026)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
por: Menon, Anjali R., et al.
Publicado: (2025)
por: Menon, Anjali R., et al.
Publicado: (2025)
The MSR-Video to Text Dataset with Clean Annotations
por: Chen, Haoran, et al.
Publicado: (2021)
por: Chen, Haoran, et al.
Publicado: (2021)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
por: Imanov, Olaf Yunus Laitinen
Publicado: (2026)
por: Imanov, Olaf Yunus Laitinen
Publicado: (2026)
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
por: Imanov, Olaf Yunus Laitinen, et al.
Publicado: (2026)
por: Imanov, Olaf Yunus Laitinen, et al.
Publicado: (2026)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
por: Kim, Heejun, et al.
Publicado: (2026)
por: Kim, Heejun, et al.
Publicado: (2026)
Ejemplares similares
-
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025) -
A Survey on Vision-Language-Action Models for Embodied AI
por: Ma, Yueen, et al.
Publicado: (2024) -
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025) -
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
por: Chen, Kewei, et al.
Publicado: (2026) -
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
por: Malikussaid, et al.
Publicado: (2026)