Valley: Video Assistant with Large Language model Enhanced abilitY
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Ruipu, Zhao, Ziwang, Yang, Min, Yang, Zheming, Qiu, Minghui, Wang, Tao, Wei, Zhongyu, Wang, Yanhao, Chen, Cen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
von: Li, Zejun, et al.
Veröffentlicht: (2024)
von: Li, Zejun, et al.
Veröffentlicht: (2024)
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
von: Wu, Ziheng, et al.
Veröffentlicht: (2025)
von: Wu, Ziheng, et al.
Veröffentlicht: (2025)
LLMvsSmall Model? Large Language Model Based Text Augmentation Enhanced Personality Detection Model
von: Hu, Linmei, et al.
Veröffentlicht: (2024)
von: Hu, Linmei, et al.
Veröffentlicht: (2024)
Valley3: Scaling Omni Foundation Models for E-commerce
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive Learning
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
FashionLOGO: Prompting Multimodal Large Language Models for Fashion Logo Embeddings
von: Wang, Zhen, et al.
Veröffentlicht: (2023)
von: Wang, Zhen, et al.
Veröffentlicht: (2023)
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
A high-order accurate moving mesh finite element method for the radial Kohn--Sham equation
von: Luo, Zheming, et al.
Veröffentlicht: (2024)
von: Luo, Zheming, et al.
Veröffentlicht: (2024)
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
von: Zhou, Guanghao, et al.
Veröffentlicht: (2025)
MatPhys: Learning Material-Aware Physics Parameters for Deformable Object Simulation from Videos
von: Yang, Yang, et al.
Veröffentlicht: (2026)
von: Yang, Yang, et al.
Veröffentlicht: (2026)
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
von: Duan, Feiyu, et al.
Veröffentlicht: (2026)
von: Duan, Feiyu, et al.
Veröffentlicht: (2026)
SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models
von: Wang, Jia, et al.
Veröffentlicht: (2025)
von: Wang, Jia, et al.
Veröffentlicht: (2025)
MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
LLMGA: Multimodal Large Language Model based Generation Assistant
von: Xia, Bin, et al.
Veröffentlicht: (2023)
von: Xia, Bin, et al.
Veröffentlicht: (2023)
Symbolic Working Memory Enhances Language Models for Complex Rule Application
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
von: Li, Yifan, et al.
Veröffentlicht: (2025)
von: Li, Yifan, et al.
Veröffentlicht: (2025)
MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline
von: Yao, Huanjin, et al.
Veröffentlicht: (2026)
von: Yao, Huanjin, et al.
Veröffentlicht: (2026)
Open-SQL Framework: Enhancing Text-to-SQL on Open-source Large Language Models
von: Chen, Xiaojun, et al.
Veröffentlicht: (2024)
von: Chen, Xiaojun, et al.
Veröffentlicht: (2024)
FedMCP: Parameter-Efficient Federated Learning with Model-Contrastive Personalization
von: Zhao, Qianyi, et al.
Veröffentlicht: (2024)
von: Zhao, Qianyi, et al.
Veröffentlicht: (2024)
SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
von: Tian, Yuhao, et al.
Veröffentlicht: (2025)
von: Tian, Yuhao, et al.
Veröffentlicht: (2025)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
von: Lu, Weiheng, et al.
Veröffentlicht: (2024)
von: Lu, Weiheng, et al.
Veröffentlicht: (2024)
Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
von: Yang, Zhongyu, et al.
Veröffentlicht: (2025)
von: Yang, Zhongyu, et al.
Veröffentlicht: (2025)
HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
von: Zhou, Ziqin, et al.
Veröffentlicht: (2025)
von: Zhou, Ziqin, et al.
Veröffentlicht: (2025)
Power-LLaVA: Large Language and Vision Assistant for Power Transmission Line Inspection
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Enhance-A-Video: Better Generated Video for Free
von: Luo, Yang, et al.
Veröffentlicht: (2025)
von: Luo, Yang, et al.
Veröffentlicht: (2025)
PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications
von: Yang, Dingkang, et al.
Veröffentlicht: (2024)
von: Yang, Dingkang, et al.
Veröffentlicht: (2024)
ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development
von: Xu, Zhenran, et al.
Veröffentlicht: (2025)
von: Xu, Zhenran, et al.
Veröffentlicht: (2025)
Towards Proactive Interactions for In-Vehicle Conversational Assistants Utilizing Large Language Models
von: Du, Huifang, et al.
Veröffentlicht: (2024)
von: Du, Huifang, et al.
Veröffentlicht: (2024)
Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
von: Hu, Yunqing, et al.
Veröffentlicht: (2025)
von: Hu, Yunqing, et al.
Veröffentlicht: (2025)
CFBenchmark: Chinese Financial Assistant Benchmark for Large Language Model
von: Lei, Yang, et al.
Veröffentlicht: (2023)
von: Lei, Yang, et al.
Veröffentlicht: (2023)
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
von: Liang, Zichen, et al.
Veröffentlicht: (2025)
von: Liang, Zichen, et al.
Veröffentlicht: (2025)
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
von: Xia, Zhongyu, et al.
Veröffentlicht: (2026)
von: Xia, Zhongyu, et al.
Veröffentlicht: (2026)
DictLLM: Harnessing Key-Value Data Structures with Large Language Models for Enhanced Medical Diagnostics
von: Guo, YiQiu, et al.
Veröffentlicht: (2024)
von: Guo, YiQiu, et al.
Veröffentlicht: (2024)
Large Language Models Leverage External Knowledge to Extend Clinical Insight Beyond Language Boundaries
von: Wu, Jiageng, et al.
Veröffentlicht: (2023)
von: Wu, Jiageng, et al.
Veröffentlicht: (2023)
Both Matter: Enhancing the Emotional Intelligence of Large Language Models without Compromising the General Intelligence
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
von: Li, Zejun, et al.
Veröffentlicht: (2024) -
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
von: Wu, Ziheng, et al.
Veröffentlicht: (2025) -
LLMvsSmall Model? Large Language Model Based Text Augmentation Enhanced Personality Detection Model
von: Hu, Linmei, et al.
Veröffentlicht: (2024) -
Valley3: Scaling Omni Foundation Models for E-commerce
von: Chen, Zeyu, et al.
Veröffentlicht: (2026) -
DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive Learning
von: Du, Mengfei, et al.
Veröffentlicht: (2024)