Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chung, Jiwan, Yoon, Janghan, Park, Junhyeong, Lee, Sangeyl, Yang, Joowon, Park, Sooyeon, Yu, Youngjae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?
von: Ka, Keummin, et al.
Veröffentlicht: (2025)
von: Ka, Keummin, et al.
Veröffentlicht: (2025)
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
von: Kim, Junhyeok, et al.
Veröffentlicht: (2025)
von: Kim, Junhyeok, et al.
Veröffentlicht: (2025)
Judge Anything: MLLM as a Judge Across Any Modality
von: Pu, Shu, et al.
Veröffentlicht: (2025)
von: Pu, Shu, et al.
Veröffentlicht: (2025)
Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation
von: Park, Jaewoo, et al.
Veröffentlicht: (2025)
von: Park, Jaewoo, et al.
Veröffentlicht: (2025)
What MLLMs Learn about When they Learn about Multimodal Reasoning
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
AnyEdit: Edit Any Knowledge Encoded in Language Models
von: Jiang, Houcheng, et al.
Veröffentlicht: (2025)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2025)
Towards Visual Text Design Transfer Across Languages
von: Choi, Yejin, et al.
Veröffentlicht: (2024)
von: Choi, Yejin, et al.
Veröffentlicht: (2024)
VAGUE: Visual Contexts Clarify Ambiguous Expressions
von: Nam, Heejeong, et al.
Veröffentlicht: (2024)
von: Nam, Heejeong, et al.
Veröffentlicht: (2024)
A11YN: aligning LLMs for accessible web UI code generation
von: Yoon, Janghan, et al.
Veröffentlicht: (2025)
von: Yoon, Janghan, et al.
Veröffentlicht: (2025)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
Multi-Vector Index Compression in Any Modality
von: Qin, Hanxiang, et al.
Veröffentlicht: (2026)
von: Qin, Hanxiang, et al.
Veröffentlicht: (2026)
ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
von: Yang, Junyao, et al.
Veröffentlicht: (2026)
von: Yang, Junyao, et al.
Veröffentlicht: (2026)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
NExT-GPT: Any-to-Any Multimodal LLM
von: Wu, Shengqiong, et al.
Veröffentlicht: (2023)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2023)
AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
von: Cheng, Dongjie, et al.
Veröffentlicht: (2026)
von: Cheng, Dongjie, et al.
Veröffentlicht: (2026)
Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text Simplification
von: Jeong, Jinhong, et al.
Veröffentlicht: (2026)
von: Jeong, Jinhong, et al.
Veröffentlicht: (2026)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
von: Luo, Run, et al.
Veröffentlicht: (2025)
von: Luo, Run, et al.
Veröffentlicht: (2025)
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
von: Lee, Seungbeen, et al.
Veröffentlicht: (2024)
von: Lee, Seungbeen, et al.
Veröffentlicht: (2024)
Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
von: Ma, Qianli, et al.
Veröffentlicht: (2025)
von: Ma, Qianli, et al.
Veröffentlicht: (2025)
The Algebra of Meaning: Why Machines Need Montague More Than Moore's Law
von: Jeong, Cheonkam, et al.
Veröffentlicht: (2025)
von: Jeong, Cheonkam, et al.
Veröffentlicht: (2025)
Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning
von: Wu, Xiangyu, et al.
Veröffentlicht: (2025)
von: Wu, Xiangyu, et al.
Veröffentlicht: (2025)
DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
von: Fu, Jiachen, et al.
Veröffentlicht: (2025)
von: Fu, Jiachen, et al.
Veröffentlicht: (2025)
Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback
von: Son, Guijin, et al.
Veröffentlicht: (2026)
von: Son, Guijin, et al.
Veröffentlicht: (2026)
Teaching Metric Distance to Discrete Autoregressive Language Models
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
MASS: Overcoming Language Bias in Image-Text Matching
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
von: Cai, Hongru, et al.
Veröffentlicht: (2026)
von: Cai, Hongru, et al.
Veröffentlicht: (2026)
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM
von: Shikhar, Sambal, et al.
Veröffentlicht: (2025)
von: Shikhar, Sambal, et al.
Veröffentlicht: (2025)
Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model
von: He, Jacqueline, et al.
Veröffentlicht: (2026)
von: He, Jacqueline, et al.
Veröffentlicht: (2026)
Toward Robust RALMs: Revealing the Impact of Imperfect Retrieval on Retrieval-Augmented Language Models
von: Park, Seong-Il, et al.
Veröffentlicht: (2024)
von: Park, Seong-Il, et al.
Veröffentlicht: (2024)
Emojinize: Enriching Any Text with Emoji Translations
von: Klein, Lars Henning, et al.
Veröffentlicht: (2024)
von: Klein, Lars Henning, et al.
Veröffentlicht: (2024)
The Emergence of Abstract Thought in Large Language Models Beyond Any Language
von: Chen, Yuxin, et al.
Veröffentlicht: (2025)
von: Chen, Yuxin, et al.
Veröffentlicht: (2025)
AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls
von: Du, Yu, et al.
Veröffentlicht: (2024)
von: Du, Yu, et al.
Veröffentlicht: (2024)
Multimodal Crystal Flow: Any-to-Any Modality Generation for Unified Crystal Modeling
von: Seong, Kiyoung, et al.
Veröffentlicht: (2026)
von: Seong, Kiyoung, et al.
Veröffentlicht: (2026)
The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2024)
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2024)
Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding
von: Tang, Yiwen, et al.
Veröffentlicht: (2024)
von: Tang, Yiwen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?
von: Ka, Keummin, et al.
Veröffentlicht: (2025) -
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
von: Kim, Junhyeok, et al.
Veröffentlicht: (2025) -
Judge Anything: MLLM as a Judge Across Any Modality
von: Pu, Shu, et al.
Veröffentlicht: (2025) -
Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation
von: Park, Jaewoo, et al.
Veröffentlicht: (2025) -
What MLLMs Learn about When they Learn about Multimodal Reasoning
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)