Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chandhok, Shivam, Yang, Qian, Manas, Oscar, Jain, Kanishk, Sigal, Leonid, Agrawal, Aishwarya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
Test-Time Consistency in Vision Language Models
von: Chou, Shih-Han, et al.
Veröffentlicht: (2025)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2025)
SceneGPT: A Language Model for 3D Scene Understanding
von: Chandhok, Shivam
Veröffentlicht: (2024)
von: Chandhok, Shivam
Veröffentlicht: (2024)
Do Vision-Language Foundational models show Robust Visual Perception?
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
Discovering Failure Modes in Vision-Language Models using RL
von: Jain, Kanishk, et al.
Veröffentlicht: (2026)
von: Jain, Kanishk, et al.
Veröffentlicht: (2026)
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
von: He, Xiangteng, et al.
Veröffentlicht: (2025)
von: He, Xiangteng, et al.
Veröffentlicht: (2025)
Assessing and Learning Alignment of Unimodal Vision and Language Models
von: Zhang, Le, et al.
Veröffentlicht: (2024)
von: Zhang, Le, et al.
Veröffentlicht: (2024)
Improving Automatic VQA Evaluation Using Large Language Models
von: Mañas, Oscar, et al.
Veröffentlicht: (2023)
von: Mañas, Oscar, et al.
Veröffentlicht: (2023)
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
von: Rahman, Tanzila, et al.
Veröffentlicht: (2024)
von: Rahman, Tanzila, et al.
Veröffentlicht: (2024)
Federated Learning with Uncertainty and Personalization via Efficient Second-order Optimization
von: Pal, Shivam, et al.
Veröffentlicht: (2024)
von: Pal, Shivam, et al.
Veröffentlicht: (2024)
Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
von: Agarwal, Aishwarya, et al.
Veröffentlicht: (2025)
von: Agarwal, Aishwarya, et al.
Veröffentlicht: (2025)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
von: Salamatian, Ali, et al.
Veröffentlicht: (2026)
von: Salamatian, Ali, et al.
Veröffentlicht: (2026)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
von: Yang, Qian, et al.
Veröffentlicht: (2026)
von: Yang, Qian, et al.
Veröffentlicht: (2026)
Controlling Multimodal LLMs via Reward-guided Decoding
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
Benchmarking Vision Language Models for Cultural Understanding
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2026)
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2026)
Aiding Medical Diagnosis through Image Synthesis and Classification
von: Choudhary, Kanishk
Veröffentlicht: (2025)
von: Choudhary, Kanishk
Veröffentlicht: (2025)
The Inductive Bottleneck: Data-Driven Emergence of Representational Sparsity in Vision Transformers
von: Awadhiya, Kanishk
Veröffentlicht: (2025)
von: Awadhiya, Kanishk
Veröffentlicht: (2025)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
von: Mañas, Oscar, et al.
Veröffentlicht: (2024)
von: Mañas, Oscar, et al.
Veröffentlicht: (2024)
PDMP: Rethinking Balanced Multimodal Learning via Performance-Dominant Modality Prioritization
von: Wei, Shicai, et al.
Veröffentlicht: (2026)
von: Wei, Shicai, et al.
Veröffentlicht: (2026)
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
von: Zhang, Le, et al.
Veröffentlicht: (2023)
von: Zhang, Le, et al.
Veröffentlicht: (2023)
RiT: Vanilla Diffusion Transformers Suffice in Representation Space
von: Zhang, Le, et al.
Veröffentlicht: (2026)
von: Zhang, Le, et al.
Veröffentlicht: (2026)
Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models
von: Xu, Bicheng, et al.
Veröffentlicht: (2024)
von: Xu, Bicheng, et al.
Veröffentlicht: (2024)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
von: Rahman, Tanzila, et al.
Veröffentlicht: (2026)
von: Rahman, Tanzila, et al.
Veröffentlicht: (2026)
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
von: Bhatt, Gaurav, et al.
Veröffentlicht: (2024)
von: Bhatt, Gaurav, et al.
Veröffentlicht: (2024)
Improving Generalization via Meta-Learning on Hard Samples
von: Jain, Nishant, et al.
Veröffentlicht: (2024)
von: Jain, Nishant, et al.
Veröffentlicht: (2024)
Enhancing Semi-Supervised Learning via Representative and Diverse Sample Selection
von: Shao, Qian, et al.
Veröffentlicht: (2024)
von: Shao, Qian, et al.
Veröffentlicht: (2024)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
von: Ahmadi, Saba, et al.
Veröffentlicht: (2023)
von: Ahmadi, Saba, et al.
Veröffentlicht: (2023)
Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling
von: Chen, Qiyuan, et al.
Veröffentlicht: (2026)
von: Chen, Qiyuan, et al.
Veröffentlicht: (2026)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
von: Chou, Shih-Han, et al.
Veröffentlicht: (2023)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2023)
Factorized Video Autoencoders for Efficient Generative Modelling
von: Suhail, Mohammed, et al.
Veröffentlicht: (2024)
von: Suhail, Mohammed, et al.
Veröffentlicht: (2024)
AGGRNet: Selective Feature Extraction and Aggregation for Enhanced Medical Image Classification
von: Makwe, Ansh, et al.
Veröffentlicht: (2025)
von: Makwe, Ansh, et al.
Veröffentlicht: (2025)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
von: Goyal, Raghav, et al.
Veröffentlicht: (2023)
von: Goyal, Raghav, et al.
Veröffentlicht: (2023)
RAVEL: Rare Concept Generation and Editing via Graph-driven Relational Guidance
von: Venkatesh, Kavana, et al.
Veröffentlicht: (2024)
von: Venkatesh, Kavana, et al.
Veröffentlicht: (2024)
Representing Animatable Avatar via Factorized Neural Fields
von: Song, Chunjin, et al.
Veröffentlicht: (2024)
von: Song, Chunjin, et al.
Veröffentlicht: (2024)
Prioritized Semantic Learning for Zero-shot Instance Navigation
von: Sun, Xinyu, et al.
Veröffentlicht: (2024)
von: Sun, Xinyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024) -
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024) -
Test-Time Consistency in Vision Language Models
von: Chou, Shih-Han, et al.
Veröffentlicht: (2025) -
SceneGPT: A Language Model for 3D Scene Understanding
von: Chandhok, Shivam
Veröffentlicht: (2024) -
Do Vision-Language Foundational models show Robust Visual Perception?
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)