Task Me Anything
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Jieyu, Huang, Weikai, Ma, Zixian, Michel, Oscar, He, Dong, Gupta, Tanmay, Ma, Wei-Chiu, Farhadi, Ali, Kembhavi, Aniruddha, Krishna, Ranjay |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
di: Ma, Zixian, et al.
Pubblicazione: (2024)
di: Ma, Zixian, et al.
Pubblicazione: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
di: Gao, Ziqi, et al.
Pubblicazione: (2024)
di: Gao, Ziqi, et al.
Pubblicazione: (2024)
Iterated Learning Improves Compositionality in Large Vision-Language Models
di: Zheng, Chenhao, et al.
Pubblicazione: (2024)
di: Zheng, Chenhao, et al.
Pubblicazione: (2024)
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
di: Duggal, Shivam, et al.
Pubblicazione: (2025)
di: Duggal, Shivam, et al.
Pubblicazione: (2025)
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
di: Yang, Yinuo, et al.
Pubblicazione: (2026)
di: Yang, Yinuo, et al.
Pubblicazione: (2026)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
di: Eftekhar, Ainaz, et al.
Pubblicazione: (2023)
di: Eftekhar, Ainaz, et al.
Pubblicazione: (2023)
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
di: Wallingford, Matthew, et al.
Pubblicazione: (2024)
di: Wallingford, Matthew, et al.
Pubblicazione: (2024)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
di: Yang, Yue, et al.
Pubblicazione: (2025)
di: Yang, Yue, et al.
Pubblicazione: (2025)
Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
di: Liu, Benlin, et al.
Pubblicazione: (2024)
di: Liu, Benlin, et al.
Pubblicazione: (2024)
Preserving Identity with Variational Score for General-purpose 3D Editing
di: Le, Duong H., et al.
Pubblicazione: (2024)
di: Le, Duong H., et al.
Pubblicazione: (2024)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
di: Gupta, Tanmay, et al.
Pubblicazione: (2026)
di: Gupta, Tanmay, et al.
Pubblicazione: (2026)
Contrastive Flow Matching
di: Stoica, George, et al.
Pubblicazione: (2025)
di: Stoica, George, et al.
Pubblicazione: (2025)
One Diffusion to Generate Them All
di: Le, Duong H., et al.
Pubblicazione: (2024)
di: Le, Duong H., et al.
Pubblicazione: (2024)
TrajTok: Learning Trajectory Tokens enables better Video Understanding
di: Zheng, Chenhao, et al.
Pubblicazione: (2026)
di: Zheng, Chenhao, et al.
Pubblicazione: (2026)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
di: Zhang, Jieyu, et al.
Pubblicazione: (2024)
di: Zhang, Jieyu, et al.
Pubblicazione: (2024)
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
di: Huang, Weikai, et al.
Pubblicazione: (2025)
di: Huang, Weikai, et al.
Pubblicazione: (2025)
Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos
di: Luo, Rundong, et al.
Pubblicazione: (2025)
di: Luo, Rundong, et al.
Pubblicazione: (2025)
The One RING: a Robotic Indoor Navigation Generalist
di: Eftekhar, Ainaz, et al.
Pubblicazione: (2024)
di: Eftekhar, Ainaz, et al.
Pubblicazione: (2024)
Synthetic Visual Genome
di: Park, Jae Sung, et al.
Pubblicazione: (2025)
di: Park, Jae Sung, et al.
Pubblicazione: (2025)
Unifying Segment Anything in Microscopy with Vision-Language Knowledge
di: Li, Manyu, et al.
Pubblicazione: (2025)
di: Li, Manyu, et al.
Pubblicazione: (2025)
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
di: Jain, Jitesh, et al.
Pubblicazione: (2025)
di: Jain, Jitesh, et al.
Pubblicazione: (2025)
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
di: Li, Linjie, et al.
Pubblicazione: (2025)
di: Li, Linjie, et al.
Pubblicazione: (2025)
LATTE: Learning to Think with Vision Specialists
di: Ma, Zixian, et al.
Pubblicazione: (2024)
di: Ma, Zixian, et al.
Pubblicazione: (2024)
FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning
di: Hu, Jiaheng, et al.
Pubblicazione: (2024)
di: Hu, Jiaheng, et al.
Pubblicazione: (2024)
Video-Based Reward Modeling for Computer-Use Agents
di: Song, Linxin, et al.
Pubblicazione: (2026)
di: Song, Linxin, et al.
Pubblicazione: (2026)
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
di: Clark, Christopher, et al.
Pubblicazione: (2026)
di: Clark, Christopher, et al.
Pubblicazione: (2026)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
di: Clark, Christopher, et al.
Pubblicazione: (2026)
di: Clark, Christopher, et al.
Pubblicazione: (2026)
SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World
di: Ehsani, Kiana, et al.
Pubblicazione: (2023)
di: Ehsani, Kiana, et al.
Pubblicazione: (2023)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
di: Maharana, Adyasha, et al.
Pubblicazione: (2023)
di: Maharana, Adyasha, et al.
Pubblicazione: (2023)
WildDet3D: Scaling Promptable 3D Detection in the Wild
di: Huang, Weikai, et al.
Pubblicazione: (2026)
di: Huang, Weikai, et al.
Pubblicazione: (2026)
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
di: He, Qijia, et al.
Pubblicazione: (2026)
di: He, Qijia, et al.
Pubblicazione: (2026)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
di: Salehi, Mohammadreza, et al.
Pubblicazione: (2024)
di: Salehi, Mohammadreza, et al.
Pubblicazione: (2024)
MIMIC: Masked Image Modeling with Image Correspondences
di: Marathe, Kalyani, et al.
Pubblicazione: (2023)
di: Marathe, Kalyani, et al.
Pubblicazione: (2023)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
di: Duan, Jiafei, et al.
Pubblicazione: (2024)
di: Duan, Jiafei, et al.
Pubblicazione: (2024)
Posterior Augmented Flow Matching
di: Stoica, George, et al.
Pubblicazione: (2026)
di: Stoica, George, et al.
Pubblicazione: (2026)
Reinforced Visual Perception with Tools
di: Zhou, Zetong, et al.
Pubblicazione: (2025)
di: Zhou, Zetong, et al.
Pubblicazione: (2025)
Multilingual Diversity Improves Vision-Language Representations
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
Seeing the Unseen: Visual Common Sense for Semantic Placement
di: Ramrakhya, Ram, et al.
Pubblicazione: (2024)
di: Ramrakhya, Ram, et al.
Pubblicazione: (2024)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
di: Gao, Ziqi, et al.
Pubblicazione: (2026)
di: Gao, Ziqi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
di: Ma, Zixian, et al.
Pubblicazione: (2024) -
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
di: Gao, Ziqi, et al.
Pubblicazione: (2024) -
Iterated Learning Improves Compositionality in Large Vision-Language Models
di: Zheng, Chenhao, et al.
Pubblicazione: (2024) -
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
di: Duggal, Shivam, et al.
Pubblicazione: (2025) -
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
di: Yang, Yinuo, et al.
Pubblicazione: (2026)