PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Blume, Ansel, Kim, Jeonghwan, Ha, Hyeonjeong, Chatikyan, Elen, Jin, Xiaomeng, Nguyen, Khanh Duy, Peng, Nanyun, Chang, Kai-Wei, Hoiem, Derek, Ji, Heng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SYNTHIA: Novel Concept Design with Affordance Composition
by: Ha, Hyeonjeong, et al.
Published: (2025)
by: Ha, Hyeonjeong, et al.
Published: (2025)
MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks
by: Ha, Hyeonjeong, et al.
Published: (2025)
by: Ha, Hyeonjeong, et al.
Published: (2025)
Search and Detect: Training-Free Long Tail Object Detection via Web-Image Retrieval
by: Sidhu, Mankeerat, et al.
Published: (2024)
by: Sidhu, Mankeerat, et al.
Published: (2024)
ARMADA: Attribute-Based Multimodal Data Augmentation
by: Jin, Xiaomeng, et al.
Published: (2024)
by: Jin, Xiaomeng, et al.
Published: (2024)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
by: Kim, Jeonghwan, et al.
Published: (2024)
by: Kim, Jeonghwan, et al.
Published: (2024)
Contrastive Visual Data Augmentation
by: Zhou, Yu, et al.
Published: (2025)
by: Zhou, Yu, et al.
Published: (2025)
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
by: Dalal, Dwip, et al.
Published: (2025)
by: Dalal, Dwip, et al.
Published: (2025)
Scientific Opinion Summarization: Paper Meta-review Generation Dataset, Methods, and Evaluation
by: Zeng, Qi, et al.
Published: (2023)
by: Zeng, Qi, et al.
Published: (2023)
Visual Program Distillation with Template-Based Augmentation
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
Region-Based Representations Revisited
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
by: Wadhawan, Rohan, et al.
Published: (2024)
by: Wadhawan, Rohan, et al.
Published: (2024)
Learning to Rank Caption Chains for Video-Text Alignment
by: Blume, Ansel, et al.
Published: (2026)
by: Blume, Ansel, et al.
Published: (2026)
RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
by: Khosla, Savya, et al.
Published: (2024)
by: Khosla, Savya, et al.
Published: (2024)
Sustainable development during economic uncertainty: What drives large construction firms to perform corporate social responsibility?
by: Minh Van Nguyen, et al.
Published: (2024)
by: Minh Van Nguyen, et al.
Published: (2024)
How to Teach Large Multimodal Models New Skills
by: Zhu, Zhen, et al.
Published: (2025)
by: Zhu, Zhen, et al.
Published: (2025)
Anytime Continual Learning for Open Vocabulary Classification
by: Zhu, Zhen, et al.
Published: (2024)
by: Zhu, Zhen, et al.
Published: (2024)
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
by: Ha, Hyeonjeong, et al.
Published: (2026)
by: Ha, Hyeonjeong, et al.
Published: (2026)
MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models
by: Ha, Hyeonjeong, et al.
Published: (2026)
by: Ha, Hyeonjeong, et al.
Published: (2026)
Language Models are Bounded Pragmatic Speakers: Understanding RLHF from a Bayesian Cognitive Modeling Perspective
by: Nguyen, Khanh
Published: (2023)
by: Nguyen, Khanh
Published: (2023)
Ensembling Portfolio Strategies for Long-Term Investments: A Distribution-Free Preference Framework for Decision-Making and Algorithms
by: Lam, Duy Khanh
Published: (2024)
by: Lam, Duy Khanh
Published: (2024)
Mean-Variance Portfolio Selection in Long-Term Investments with Unknown Distribution: Online Estimation, Risk Aversion under Ambiguity, and Universality of Algorithms
by: Lam, Duy Khanh
Published: (2024)
by: Lam, Duy Khanh
Published: (2024)
Beating the Best Constant Rebalancing Portfolio in Long-Term Investment: A Generalization of the Kelly Criterion and Universal Learning Algorithm for Markets with Serial Dependence
by: Lam, Duy Khanh
Published: (2025)
by: Lam, Duy Khanh
Published: (2025)
Sequential Portfolio Selection under Latent Side Information-Dependence Structure: Optimality and Universal Learning Algorithms
by: Lam, Duy Khanh
Published: (2025)
by: Lam, Duy Khanh
Published: (2025)
Perception-Aware Policy Optimization for Multimodal Reasoning
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
by: Hu, Wenbo, et al.
Published: (2026)
by: Hu, Wenbo, et al.
Published: (2026)
Continual Learning in Open-vocabulary Classification with Complementary Memory Systems
by: Zhu, Zhen, et al.
Published: (2023)
by: Zhu, Zhen, et al.
Published: (2023)
Quantum States Seen by a Probe: Partial Trace Over a Region of Space
by: Ansel, Quentin
Published: (2024)
by: Ansel, Quentin
Published: (2024)
Emergent gravity from the correlation of spin-$\tfrac{1}{2}$ systems coupled with a scalar field
by: Ansel, Quentin
Published: (2024)
by: Ansel, Quentin
Published: (2024)
Perspective of a pre‐symptomatic individual with an FTD‐causing MAPT gene variant
by: Ansel Dow
Published: (2025)
by: Ansel Dow
Published: (2025)
CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing
by: Qian, Cheng, et al.
Published: (2026)
by: Qian, Cheng, et al.
Published: (2026)
Sustainable Supply Chain Practices and the Cost of Equity Capital in Emerging Markets: Do Blockholder Ownership and Geopolitical Risk Matter?
by: Van Ha Nguyen, et al.
Published: (2026)
by: Van Ha Nguyen, et al.
Published: (2026)
Automatic Textual Normalization for Hate Speech Detection
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2023)
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2023)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
by: Zhan, Qiusi, et al.
Published: (2025)
by: Zhan, Qiusi, et al.
Published: (2025)
VirDA: Reusing Backbone for Unsupervised Domain Adaptation with Visual Reprogramming
by: Nguyen, Duy, et al.
Published: (2025)
by: Nguyen, Duy, et al.
Published: (2025)
Federated Document Visual Question Answering: A Pilot Study
by: Nguyen, Khanh, et al.
Published: (2024)
by: Nguyen, Khanh, et al.
Published: (2024)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
by: Wu, Te-Lin, et al.
Published: (2021)
by: Wu, Te-Lin, et al.
Published: (2021)
Domain Generalization through Spatial Relation Induction over Visual Primitives
by: Nguyen, Dat, et al.
Published: (2026)
by: Nguyen, Dat, et al.
Published: (2026)
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
by: Khosla, Savya, et al.
Published: (2026)
by: Khosla, Savya, et al.
Published: (2026)
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
by: Khosla, Savya, et al.
Published: (2025)
by: Khosla, Savya, et al.
Published: (2025)
Similar Items
-
SYNTHIA: Novel Concept Design with Affordance Composition
by: Ha, Hyeonjeong, et al.
Published: (2025) -
MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks
by: Ha, Hyeonjeong, et al.
Published: (2025) -
Search and Detect: Training-Free Long Tail Object Detection via Web-Image Retrieval
by: Sidhu, Mankeerat, et al.
Published: (2024) -
ARMADA: Attribute-Based Multimodal Data Augmentation
by: Jin, Xiaomeng, et al.
Published: (2024) -
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
by: Kim, Jeonghwan, et al.
Published: (2024)