EVA-02: A Visual Representation for Neon Genesis
Fuente:
arXiv
Guardado en:
| Autores principales: | Fang, Yuxin, Sun, Quan, Wang, Xinggang, Huang, Tiejun, Wang, Xinlong, Cao, Yue |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
por: Sun, Quan, et al.
Publicado: (2024)
por: Sun, Quan, et al.
Publicado: (2024)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
por: Zhu, Lianghui, et al.
Publicado: (2024)
por: Zhu, Lianghui, et al.
Publicado: (2024)
CapsFusion: Rethinking Image-Text Data at Scale
por: Yu, Qiying, et al.
Publicado: (2023)
por: Yu, Qiying, et al.
Publicado: (2023)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
por: Zhang, Yaolun, et al.
Publicado: (2026)
por: Zhang, Yaolun, et al.
Publicado: (2026)
Exploring the Design Space of Visual Context Representation in Video MLLMs
por: Du, Yifan, et al.
Publicado: (2024)
por: Du, Yifan, et al.
Publicado: (2024)
EVA-X: A Foundation Model for General Chest X-ray Analysis with Self-supervised Learning
por: Yao, Jingfeng, et al.
Publicado: (2024)
por: Yao, Jingfeng, et al.
Publicado: (2024)
Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing
por: Yuan, Fan, et al.
Publicado: (2025)
por: Yuan, Fan, et al.
Publicado: (2025)
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
Visual Representations inside the Language Model
por: Liu, Benlin, et al.
Publicado: (2025)
por: Liu, Benlin, et al.
Publicado: (2025)
Emu: Generative Pretraining in Multimodality
por: Sun, Quan, et al.
Publicado: (2023)
por: Sun, Quan, et al.
Publicado: (2023)
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
por: Wen, Yuxin, et al.
Publicado: (2024)
por: Wen, Yuxin, et al.
Publicado: (2024)
Emergent Visual-Semantic Hierarchies in Image-Text Representations
por: Alper, Morris, et al.
Publicado: (2024)
por: Alper, Morris, et al.
Publicado: (2024)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
por: Wang, Weiyun, et al.
Publicado: (2025)
por: Wang, Weiyun, et al.
Publicado: (2025)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
por: Zheng, Duo, et al.
Publicado: (2024)
por: Zheng, Duo, et al.
Publicado: (2024)
Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning
por: Li, Mingcheng, et al.
Publicado: (2024)
por: Li, Mingcheng, et al.
Publicado: (2024)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
por: Xu, Xiaohao, et al.
Publicado: (2024)
por: Xu, Xiaohao, et al.
Publicado: (2024)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
por: Li, Zhuowan, et al.
Publicado: (2022)
por: Li, Zhuowan, et al.
Publicado: (2022)
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
por: Bi, Jinhe, et al.
Publicado: (2024)
por: Bi, Jinhe, et al.
Publicado: (2024)
Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models
por: Cheng, Hao, et al.
Publicado: (2025)
por: Cheng, Hao, et al.
Publicado: (2025)
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
por: Guo, Xiangyu, et al.
Publicado: (2025)
por: Guo, Xiangyu, et al.
Publicado: (2025)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
por: Jiang, Houcheng, et al.
Publicado: (2026)
por: Jiang, Houcheng, et al.
Publicado: (2026)
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
por: Wang, Chenglong, et al.
Publicado: (2024)
por: Wang, Chenglong, et al.
Publicado: (2024)
Latent Visual Reasoning
por: Li, Bangzheng, et al.
Publicado: (2025)
por: Li, Bangzheng, et al.
Publicado: (2025)
Towards Scalable Pre-training of Visual Tokenizers for Generation
por: Yao, Jingfeng, et al.
Publicado: (2025)
por: Yao, Jingfeng, et al.
Publicado: (2025)
Generative Multimodal Models are In-Context Learners
por: Sun, Quan, et al.
Publicado: (2023)
por: Sun, Quan, et al.
Publicado: (2023)
DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation
por: Wang, Jiapeng, et al.
Publicado: (2024)
por: Wang, Jiapeng, et al.
Publicado: (2024)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
por: Liu, Fanfan, et al.
Publicado: (2024)
por: Liu, Fanfan, et al.
Publicado: (2024)
Refining Skewed Perceptions in Vision-Language Contrastive Models through Visual Representations
por: Dai, Haocheng, et al.
Publicado: (2024)
por: Dai, Haocheng, et al.
Publicado: (2024)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
por: Cheng, Zihui, et al.
Publicado: (2025)
por: Cheng, Zihui, et al.
Publicado: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
por: Du, Yifan, et al.
Publicado: (2023)
por: Du, Yifan, et al.
Publicado: (2023)
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
por: Xing, Long, et al.
Publicado: (2024)
por: Xing, Long, et al.
Publicado: (2024)
Superpixel Semantics Representation and Pre-training for Vision-Language Task
por: Zhang, Siyu, et al.
Publicado: (2023)
por: Zhang, Siyu, et al.
Publicado: (2023)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
por: Fang, I-Sheng, et al.
Publicado: (2025)
por: Fang, I-Sheng, et al.
Publicado: (2025)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
por: Jian, Pu, et al.
Publicado: (2025)
por: Jian, Pu, et al.
Publicado: (2025)
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
por: Mi, Hongze, et al.
Publicado: (2024)
por: Mi, Hongze, et al.
Publicado: (2024)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
por: Shi, Weikang, et al.
Publicado: (2025)
por: Shi, Weikang, et al.
Publicado: (2025)
Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education
por: Wang, Junling, et al.
Publicado: (2026)
por: Wang, Junling, et al.
Publicado: (2026)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
por: Hua, Jiacheng, et al.
Publicado: (2026)
por: Hua, Jiacheng, et al.
Publicado: (2026)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
por: Song, Wei, et al.
Publicado: (2025)
por: Song, Wei, et al.
Publicado: (2025)
OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning
por: Liu, Yuecheng, et al.
Publicado: (2025)
por: Liu, Yuecheng, et al.
Publicado: (2025)
Ejemplares similares
-
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
por: Sun, Quan, et al.
Publicado: (2024) -
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
por: Zhu, Lianghui, et al.
Publicado: (2024) -
CapsFusion: Rethinking Image-Text Data at Scale
por: Yu, Qiying, et al.
Publicado: (2023) -
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
por: Zhang, Yaolun, et al.
Publicado: (2026) -
Exploring the Design Space of Visual Context Representation in Video MLLMs
por: Du, Yifan, et al.
Publicado: (2024)