Activation Reward Models for Few-Shot Model Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chai, Tianning, Mitra, Chancharik, Huang, Brandon, Gare, Gautam Rajendrakumar, Lin, Zhiqiu, Arbelle, Assaf, Karlinsky, Leonid, Feris, Rogerio, Darrell, Trevor, Ramanan, Deva, Herzig, Roei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2024)
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2024)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2025)
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2025)
Latent Implicit Visual Reasoning
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection
von: Gare, Gautam Rajendrakumar, et al.
Veröffentlicht: (2026)
von: Gare, Gautam Rajendrakumar, et al.
Veröffentlicht: (2026)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
von: Huang, Irene, et al.
Veröffentlicht: (2024)
von: Huang, Irene, et al.
Veröffentlicht: (2024)
Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
Revisiting Few-Shot Object Detection with Vision-Language Models
von: Madan, Anish, et al.
Veröffentlicht: (2023)
von: Madan, Anish, et al.
Veröffentlicht: (2023)
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
Improving Model's Interpretability and Reliability using Biomarkers
von: Gare, Gautam Rajendrakumar, et al.
Veröffentlicht: (2024)
von: Gare, Gautam Rajendrakumar, et al.
Veröffentlicht: (2024)
NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning
von: Schwartz, Eli, et al.
Veröffentlicht: (2024)
von: Schwartz, Eli, et al.
Veröffentlicht: (2024)
State-Space Large Audio Language Models
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
Recursive Visual Programming
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
CAMELoT: Towards Large Language Models with Training-Free Consolidated Associative Memory
von: He, Zexue, et al.
Veröffentlicht: (2024)
von: He, Zexue, et al.
Veröffentlicht: (2024)
Towards Multimodal In-Context Learning for Vision & Language Models
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
$\texttt{BATCLIP}$: Bimodal Online Test-Time Adaptation for CLIP
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2024)
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2024)
MAEDAY: MAE for few and zero shot AnomalY-Detection
von: Schwartz, Eli, et al.
Veröffentlicht: (2022)
von: Schwartz, Eli, et al.
Veröffentlicht: (2022)
Revisiting the Role of Language Priors in Vision-Language Models
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
Teaching VLMs to Localize Specific Objects from In-context Examples
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
von: Hsieh, Wen-Han, et al.
Veröffentlicht: (2025)
von: Hsieh, Wen-Han, et al.
Veröffentlicht: (2025)
In-Context Learning Enables Robot Action Prediction in LLMs
von: Yin, Yida, et al.
Veröffentlicht: (2024)
von: Yin, Yida, et al.
Veröffentlicht: (2024)
Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
Pre-training Auto-regressive Robotic Models with 4D Representations
von: Niu, Dantong, et al.
Veröffentlicht: (2025)
von: Niu, Dantong, et al.
Veröffentlicht: (2025)
Language Models as Black-Box Optimizers for Vision-Language Models
von: Liu, Shihong, et al.
Veröffentlicht: (2023)
von: Liu, Shihong, et al.
Veröffentlicht: (2023)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
$\textit{Trans-LoRA}$: towards data-free Transferable Parameter Efficient Finetuning
von: Wang, Runqian, et al.
Veröffentlicht: (2024)
von: Wang, Runqian, et al.
Veröffentlicht: (2024)
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
von: Kondic, Jovana, et al.
Veröffentlicht: (2025)
von: Kondic, Jovana, et al.
Veröffentlicht: (2025)
Self-Specialization: Uncovering Latent Expertise within Large Language Models
von: Kang, Junmo, et al.
Veröffentlicht: (2023)
von: Kang, Junmo, et al.
Veröffentlicht: (2023)
Comparison Visual Instruction Tuning
von: Lin, Wei, et al.
Veröffentlicht: (2024)
von: Lin, Wei, et al.
Veröffentlicht: (2024)
Adaptive Memory Replay for Continual Learning
von: Smith, James Seale, et al.
Veröffentlicht: (2024)
von: Smith, James Seale, et al.
Veröffentlicht: (2024)
Large Scale Generative AI Text Applied to Sports and Music
von: Baughman, Aaron, et al.
Veröffentlicht: (2024)
von: Baughman, Aaron, et al.
Veröffentlicht: (2024)
The Neglected Tails in Vision-Language Models
von: Parashar, Shubham, et al.
Veröffentlicht: (2024)
von: Parashar, Shubham, et al.
Veröffentlicht: (2024)
Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
von: Kang, Junmo, et al.
Veröffentlicht: (2024)
von: Kang, Junmo, et al.
Veröffentlicht: (2024)
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024) -
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
von: Huang, Brandon, et al.
Veröffentlicht: (2024) -
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023) -
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2024) -
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)