Mull-Tokens: Modality-Agnostic Latent Thinking
Fuente:
arXiv
Salvato in:
| Autori principali: | Ray, Arijit, Abdelkader, Ahmed, Mao, Chengzhi, Plummer, Bryan A., Saenko, Kate, Krishna, Ranjay, Guibas, Leonidas, Chu, Wen-Sheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tell Me What's Next: Textual Foresight for Generic UI Representations
di: Burns, Andrea, et al.
Pubblicazione: (2024)
di: Burns, Andrea, et al.
Pubblicazione: (2024)
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
di: Qraitem, Maan, et al.
Pubblicazione: (2023)
di: Qraitem, Maan, et al.
Pubblicazione: (2023)
SLANT: Spurious Logo ANalysis Toolkit
di: Qraitem, Maan, et al.
Pubblicazione: (2024)
di: Qraitem, Maan, et al.
Pubblicazione: (2024)
Web Artifact Attacks Disrupt Vision Language Models
di: Qraitem, Maan, et al.
Pubblicazione: (2025)
di: Qraitem, Maan, et al.
Pubblicazione: (2025)
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
di: Ray, Arijit, et al.
Pubblicazione: (2024)
di: Ray, Arijit, et al.
Pubblicazione: (2024)
OP-LoRA: The Blessing of Dimensionality
di: Teterwak, Piotr, et al.
Pubblicazione: (2024)
di: Teterwak, Piotr, et al.
Pubblicazione: (2024)
Is Large-Scale Pretraining the Secret to Good Domain Generalization?
di: Teterwak, Piotr, et al.
Pubblicazione: (2024)
di: Teterwak, Piotr, et al.
Pubblicazione: (2024)
ERM++: An Improved Baseline for Domain Generalization
di: Teterwak, Piotr, et al.
Pubblicazione: (2023)
di: Teterwak, Piotr, et al.
Pubblicazione: (2023)
CLAMP: Contrastive LAnguage Model Prompt-tuning
di: Teterwak, Piotr, et al.
Pubblicazione: (2023)
di: Teterwak, Piotr, et al.
Pubblicazione: (2023)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
di: Qraitem, Maan, et al.
Pubblicazione: (2024)
di: Qraitem, Maan, et al.
Pubblicazione: (2024)
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
di: Brown, Ellis, et al.
Pubblicazione: (2025)
di: Brown, Ellis, et al.
Pubblicazione: (2025)
Koala: Key frame-conditioned long video-LLM
di: Tan, Reuben, et al.
Pubblicazione: (2024)
di: Tan, Reuben, et al.
Pubblicazione: (2024)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
di: Liu, Aoming, et al.
Pubblicazione: (2025)
di: Liu, Aoming, et al.
Pubblicazione: (2025)
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
di: Petsiuk, Vitali, et al.
Pubblicazione: (2024)
di: Petsiuk, Vitali, et al.
Pubblicazione: (2024)
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models
di: Huang, Ian, et al.
Pubblicazione: (2024)
di: Huang, Ian, et al.
Pubblicazione: (2024)
Marginalized Bundle Adjustment: Multi-View Camera Pose from Monocular Depth Estimates
di: Zhu, Shengjie, et al.
Pubblicazione: (2026)
di: Zhu, Shengjie, et al.
Pubblicazione: (2026)
Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models
di: Chu, Wen-Hsuan, et al.
Pubblicazione: (2023)
di: Chu, Wen-Hsuan, et al.
Pubblicazione: (2023)
Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
di: Liu, Qihao, et al.
Pubblicazione: (2025)
di: Liu, Qihao, et al.
Pubblicazione: (2025)
Modality Agnostic Efficient Long Range Encoder
di: Parag, Toufiq, et al.
Pubblicazione: (2025)
di: Parag, Toufiq, et al.
Pubblicazione: (2025)
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
di: Mishra, Samarth, et al.
Pubblicazione: (2025)
di: Mishra, Samarth, et al.
Pubblicazione: (2025)
OCH3R: Object-Centric Holistic 3D Reconstruction
di: Du, Yi, et al.
Pubblicazione: (2026)
di: Du, Yi, et al.
Pubblicazione: (2026)
InfoGaussian: Structure-Aware Dynamic Gaussians through Lightweight Information Shaping
di: Zhang, Yunchao, et al.
Pubblicazione: (2024)
di: Zhang, Yunchao, et al.
Pubblicazione: (2024)
PASTA: Controllable Part-Aware Shape Generation with Autoregressive Transformers
di: Li, Songlin, et al.
Pubblicazione: (2024)
di: Li, Songlin, et al.
Pubblicazione: (2024)
LNL+K: Enhancing Learning with Noisy Labels Through Noise Source Knowledge Integration
di: Wang, Siqi, et al.
Pubblicazione: (2023)
di: Wang, Siqi, et al.
Pubblicazione: (2023)
NeRF Revisited: Fixing Quadrature Instability in Volume Rendering
di: Uy, Mikaela Angelina, et al.
Pubblicazione: (2023)
di: Uy, Mikaela Angelina, et al.
Pubblicazione: (2023)
Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruning
di: Qin, Wenda, et al.
Pubblicazione: (2025)
di: Qin, Wenda, et al.
Pubblicazione: (2025)
RefTok: Reference-Based Tokenization for Video Generation
di: Fan, Xiang, et al.
Pubblicazione: (2025)
di: Fan, Xiang, et al.
Pubblicazione: (2025)
Refining Pre-Trained Motion Models
di: Sun, Xinglong, et al.
Pubblicazione: (2024)
di: Sun, Xinglong, et al.
Pubblicazione: (2024)
Support-Set Context Matters for Bongard Problems
di: Raghuraman, Nikhil, et al.
Pubblicazione: (2023)
di: Raghuraman, Nikhil, et al.
Pubblicazione: (2023)
SuperDec: 3D Scene Decomposition with Superquadric Primitives
di: Fedele, Elisabetta, et al.
Pubblicazione: (2025)
di: Fedele, Elisabetta, et al.
Pubblicazione: (2025)
Robust Human Registration with Body Part Segmentation on Noisy Point Clouds
di: Lascheit, Kai, et al.
Pubblicazione: (2025)
di: Lascheit, Kai, et al.
Pubblicazione: (2025)
MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
di: Lei, Jiahui, et al.
Pubblicazione: (2025)
di: Lei, Jiahui, et al.
Pubblicazione: (2025)
Dynamic Reflections: Probing Video Representations with Text Alignment
di: Zhu, Tyler, et al.
Pubblicazione: (2025)
di: Zhu, Tyler, et al.
Pubblicazione: (2025)
SplatTalk: 3D VQA with Gaussian Splatting
di: Thai, Anh, et al.
Pubblicazione: (2025)
di: Thai, Anh, et al.
Pubblicazione: (2025)
Asymmetric Flow Models
di: Chen, Hansheng, et al.
Pubblicazione: (2026)
di: Chen, Hansheng, et al.
Pubblicazione: (2026)
Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning
di: You, Yang, et al.
Pubblicazione: (2024)
di: You, Yang, et al.
Pubblicazione: (2024)
Zero-Shot Image Feature Consensus with Deep Functional Maps
di: Cheng, Xinle, et al.
Pubblicazione: (2024)
di: Cheng, Xinle, et al.
Pubblicazione: (2024)
RECAST: Reparameterized, Compact weight Adaptation for Sequential Tasks
di: Tasnim, Nazia, et al.
Pubblicazione: (2024)
di: Tasnim, Nazia, et al.
Pubblicazione: (2024)
Enhancing Feature Diversity Boosts Channel-Adaptive Vision Transformers
di: Pham, Chau, et al.
Pubblicazione: (2024)
di: Pham, Chau, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Tell Me What's Next: Textual Foresight for Generic UI Representations
di: Burns, Andrea, et al.
Pubblicazione: (2024) -
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
di: Qraitem, Maan, et al.
Pubblicazione: (2023) -
SLANT: Spurious Logo ANalysis Toolkit
di: Qraitem, Maan, et al.
Pubblicazione: (2024) -
Web Artifact Attacks Disrupt Vision Language Models
di: Qraitem, Maan, et al.
Pubblicazione: (2025) -
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
di: Ray, Arijit, et al.
Pubblicazione: (2024)