ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Xue, Le, Yu, Ning, Zhang, Shu, Panagopoulou, Artemis, Li, Junnan, Martín-Martín, Roberto, Wu, Jiajun, Xiong, Caiming, Xu, Ran, Niebles, Juan Carlos, Savarese, Silvio |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
by: Panagopoulou, Artemis, et al.
Published: (2023)
by: Panagopoulou, Artemis, et al.
Published: (2023)
ViUniT: Visual Unit Tests for More Robust Visual Programming
by: Panagopoulou, Artemis, et al.
Published: (2024)
by: Panagopoulou, Artemis, et al.
Published: (2024)
Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
by: Panagopoulou, Artemis, et al.
Published: (2025)
by: Panagopoulou, Artemis, et al.
Published: (2025)
Hierarchical Point Attention for Indoor 3D Object Detection
by: Shu, Manli, et al.
Published: (2023)
by: Shu, Manli, et al.
Published: (2023)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
by: Wang, Ziyang, et al.
Published: (2025)
by: Wang, Ziyang, et al.
Published: (2025)
Future Optical Flow Prediction Improves Robot Control & Video Generation
by: Ranasinghe, Kanchana, et al.
Published: (2026)
by: Ranasinghe, Kanchana, et al.
Published: (2026)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
by: Zhou, Honglu, et al.
Published: (2025)
by: Zhou, Honglu, et al.
Published: (2025)
MapTrace: Scalable Data Generation for Route Tracing on Maps
by: Panagopoulou, Artemis, et al.
Published: (2025)
by: Panagopoulou, Artemis, et al.
Published: (2025)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
by: Ryoo, Michael S., et al.
Published: (2024)
by: Ryoo, Michael S., et al.
Published: (2024)
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
Evaluating Vision-Language Models on Bistable Images
by: Panagopoulou, Artemis, et al.
Published: (2024)
by: Panagopoulou, Artemis, et al.
Published: (2024)
Enabling High Data Throughput Reinforcement Learning on GPUs: A Domain Agnostic Framework for Data-Driven Scientific Research
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
Causal Layering via Conditional Entropy
by: Feigenbaum, Itai, et al.
Published: (2024)
by: Feigenbaum, Itai, et al.
Published: (2024)
Editing Arbitrary Propositions in LLMs without Subject Labels
by: Feigenbaum, Itai, et al.
Published: (2024)
by: Feigenbaum, Itai, et al.
Published: (2024)
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
Shared Imagination: LLMs Hallucinate Alike
by: Zhou, Yilun, et al.
Published: (2024)
by: Zhou, Yilun, et al.
Published: (2024)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Artificial intelligence and democracy: Towards digital authoritarianism or a democratic upgrade?
by: Panagopoulou, Fereniki
Published: (2025)
by: Panagopoulou, Fereniki
Published: (2025)
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
by: Fang, Yu, et al.
Published: (2025)
by: Fang, Yu, et al.
Published: (2025)
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
by: Li, Jierui, et al.
Published: (2024)
by: Li, Jierui, et al.
Published: (2024)
W&D:Scaling Parallel Tool Calling for Efficient Deep Research Agents
by: Lin, Xiaoqiang, et al.
Published: (2026)
by: Lin, Xiaoqiang, et al.
Published: (2026)
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
by: Pang, Bo, et al.
Published: (2025)
by: Pang, Bo, et al.
Published: (2025)
MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
by: Awadalla, Anas, et al.
Published: (2024)
by: Awadalla, Anas, et al.
Published: (2024)
EVALUATINGTHEEFFECTIVENESSOFE LSS AND ULIP PROGRAMS FOR TAXPREVENTION
by: Dr.S CHANDBASHA
Published: (2026)
by: Dr.S CHANDBASHA
Published: (2026)
LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer
by: Yu, Ning, et al.
Published: (2022)
by: Yu, Ning, et al.
Published: (2022)
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
WALT: Web Agents that Learn Tools
by: Prabhu, Viraj, et al.
Published: (2025)
by: Prabhu, Viraj, et al.
Published: (2025)
Text2Data: Low-Resource Data Generation with Textual Control
by: Wang, Shiyu, et al.
Published: (2024)
by: Wang, Shiyu, et al.
Published: (2024)
Scalable Chain of Thoughts via Elastic Reasoning
by: Xu, Yuhui, et al.
Published: (2025)
by: Xu, Yuhui, et al.
Published: (2025)
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
by: Chen, Jiuhai, et al.
Published: (2025)
by: Chen, Jiuhai, et al.
Published: (2025)
LATTE: Learning to Think with Vision Specialists
by: Ma, Zixian, et al.
Published: (2024)
by: Ma, Zixian, et al.
Published: (2024)
Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization
by: Yao, Weiran, et al.
Published: (2023)
by: Yao, Weiran, et al.
Published: (2023)
REX: Rapid Exploration and eXploitation for AI Agents
by: Murthy, Rithesh, et al.
Published: (2023)
by: Murthy, Rithesh, et al.
Published: (2023)
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
by: Pang, Bo, et al.
Published: (2025)
by: Pang, Bo, et al.
Published: (2025)
Unified Training of Universal Time Series Forecasting Transformers
by: Woo, Gerald, et al.
Published: (2024)
by: Woo, Gerald, et al.
Published: (2024)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
by: Nguyen, Xuan-Phi, et al.
Published: (2025)
by: Nguyen, Xuan-Phi, et al.
Published: (2025)
BLIP3o-NEXT: Next Frontier of Native Image Generation
by: Chen, Jiuhai, et al.
Published: (2025)
by: Chen, Jiuhai, et al.
Published: (2025)
Empowering Time Series Analysis with Synthetic Data: A Survey and Outlook in the Era of Foundation Models
by: Liu, Xu, et al.
Published: (2025)
by: Liu, Xu, et al.
Published: (2025)
Moirai 2.0: When Less Is More for Time Series Forecasting
by: Liu, Chenghao, et al.
Published: (2025)
by: Liu, Chenghao, et al.
Published: (2025)
Similar Items
-
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
by: Panagopoulou, Artemis, et al.
Published: (2023) -
ViUniT: Visual Unit Tests for More Robust Visual Programming
by: Panagopoulou, Artemis, et al.
Published: (2024) -
Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
by: Panagopoulou, Artemis, et al.
Published: (2025) -
Hierarchical Point Attention for Indoor 3D Object Detection
by: Shu, Manli, et al.
Published: (2023) -
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
by: Wang, Ziyang, et al.
Published: (2025)