Enregistré dans:
| Auteurs principaux: | Liang, Chao, Ma, Fan, Zhu, Linchao, Deng, Yingying, Yang, Yi |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2402.00627 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
3DID: Direct 3D Inverse Design for Aerodynamics with Physics-Aware Optimization
par: Hao, Yuze, et autres
Publié: (2025)
par: Hao, Yuze, et autres
Publié: (2025)
From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment
par: Suo, Yucheng, et autres
Publié: (2025)
par: Suo, Yucheng, et autres
Publié: (2025)
OpenMoCap: Rethinking Optical Motion Capture under Real-world Occlusion
par: Qian, Chen, et autres
Publié: (2025)
par: Qian, Chen, et autres
Publié: (2025)
Combating Label Noise With A General Surrogate Model For Sample Selection
par: Liang, Chao, et autres
Publié: (2023)
par: Liang, Chao, et autres
Publié: (2023)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
par: Fan, Tiehan, et autres
Publié: (2024)
par: Fan, Tiehan, et autres
Publié: (2024)
Computation-Efficient and Recognition-Friendly 3D Point Cloud Privacy Protection
par: Ma, Haotian, et autres
Publié: (2025)
par: Ma, Haotian, et autres
Publié: (2025)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
par: Yuan, Huaying, et autres
Publié: (2025)
par: Yuan, Huaying, et autres
Publié: (2025)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
par: Chen, Yuyan, et autres
Publié: (2024)
par: Chen, Yuyan, et autres
Publié: (2024)
Knowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval
par: Suo, Yucheng, et autres
Publié: (2024)
par: Suo, Yucheng, et autres
Publié: (2024)
Latent-Info and Low-Dimensional Learning for Human Mesh Recovery and Parallel Optimization
par: Zhang, Xiang, et autres
Publié: (2025)
par: Zhang, Xiang, et autres
Publié: (2025)
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
par: Cao, Zhuo, et autres
Publié: (2025)
par: Cao, Zhuo, et autres
Publié: (2025)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
par: Li, Yuying, et autres
Publié: (2025)
par: Li, Yuying, et autres
Publié: (2025)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
par: Liu, Xiaolin, et autres
Publié: (2026)
par: Liu, Xiaolin, et autres
Publié: (2026)
When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning
par: Luo, Junwei, et autres
Publié: (2025)
par: Luo, Junwei, et autres
Publié: (2025)
Who Can We Trust? Scope-Aware Video Moment Retrieval with Multi-Agent Conflict
par: Wu, Chaochen, et autres
Publié: (2025)
par: Wu, Chaochen, et autres
Publié: (2025)
MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
par: Park, Seojeong, et autres
Publié: (2024)
par: Park, Seojeong, et autres
Publié: (2024)
CDUPatch: Color-Driven Universal Adversarial Patch Attack for Dual-Modal Visible-Infrared Detectors
par: Long, Jiahuan, et autres
Publié: (2025)
par: Long, Jiahuan, et autres
Publié: (2025)
Theorem-Validated Reverse Chain-of-Thought Problem Generation for Geometric Reasoning
par: Deng, Linger, et autres
Publié: (2024)
par: Deng, Linger, et autres
Publié: (2024)
Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection
par: Zheng, Haowen, et autres
Publié: (2025)
par: Zheng, Haowen, et autres
Publié: (2025)
Accelerating Video Generation Inference with Sequential-Parallel 3D Positional Encoding Using a Global Time Index
par: Yuan, Chao, et autres
Publié: (2026)
par: Yuan, Chao, et autres
Publié: (2026)
EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
par: Yang, Xiangpeng, et autres
Publié: (2024)
par: Yang, Xiangpeng, et autres
Publié: (2024)
VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing
par: Yang, Xiangpeng, et autres
Publié: (2025)
par: Yang, Xiangpeng, et autres
Publié: (2025)
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
par: Liang, Baoyu, et autres
Publié: (2025)
par: Liang, Baoyu, et autres
Publié: (2025)
Multi-scale Temporal Fusion Transformer for Incomplete Vehicle Trajectory Prediction
par: Liu, Zhanwen, et autres
Publié: (2024)
par: Liu, Zhanwen, et autres
Publié: (2024)
InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
par: Wang, Zhenzhi, et autres
Publié: (2025)
par: Wang, Zhenzhi, et autres
Publié: (2025)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
par: Xing, Long, et autres
Publié: (2025)
par: Xing, Long, et autres
Publié: (2025)
Hulk: A Universal Knowledge Translator for Human-Centric Tasks
par: Wang, Yizhou, et autres
Publié: (2023)
par: Wang, Yizhou, et autres
Publié: (2023)
DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
par: Li, Binbin, et autres
Publié: (2025)
par: Li, Binbin, et autres
Publié: (2025)
AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
par: Chen, Zigeng, et autres
Publié: (2024)
par: Chen, Zigeng, et autres
Publié: (2024)
ReGenNet: Towards Human Action-Reaction Synthesis
par: Xu, Liang, et autres
Publié: (2024)
par: Xu, Liang, et autres
Publié: (2024)
IE-Bench: Advancing the Measurement of Text-Driven Image Editing for Human Perception Alignment
par: Sun, Shangkun, et autres
Publié: (2025)
par: Sun, Shangkun, et autres
Publié: (2025)
Mitigating Vanishing Activations in Deep CapsNets Using Channel Pruning
par: Sahu, Siddharth, et autres
Publié: (2024)
par: Sahu, Siddharth, et autres
Publié: (2024)
FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
par: Wang, Yuxuan, et autres
Publié: (2025)
par: Wang, Yuxuan, et autres
Publié: (2025)
Text-guided 3D Human Motion Generation with Keyframe-based Parallel Skip Transformer
par: Geng, Zichen, et autres
Publié: (2024)
par: Geng, Zichen, et autres
Publié: (2024)
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation
par: Zhu, Yuanbing, et autres
Publié: (2024)
par: Zhu, Yuanbing, et autres
Publié: (2024)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
par: Chen, Houlun, et autres
Publié: (2024)
par: Chen, Houlun, et autres
Publié: (2024)
Noise-Tolerant Hybrid Prototypical Learning with Noisy Web Data
par: Liang, Chao, et autres
Publié: (2025)
par: Liang, Chao, et autres
Publié: (2025)
PRIME: Protect Your Videos From Malicious Editing
par: Li, Guanlin, et autres
Publié: (2024)
par: Li, Guanlin, et autres
Publié: (2024)
A Unified Perspective for Loss-Oriented Imbalanced Learning via Localization
par: Wang, Zitai, et autres
Publié: (2023)
par: Wang, Zitai, et autres
Publié: (2023)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
par: Tu, Yunbin, et autres
Publié: (2024)
par: Tu, Yunbin, et autres
Publié: (2024)
Documents similaires
-
3DID: Direct 3D Inverse Design for Aerodynamics with Physics-Aware Optimization
par: Hao, Yuze, et autres
Publié: (2025) -
From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment
par: Suo, Yucheng, et autres
Publié: (2025) -
OpenMoCap: Rethinking Optical Motion Capture under Real-world Occlusion
par: Qian, Chen, et autres
Publié: (2025) -
Combating Label Noise With A General Surrogate Model For Sample Selection
par: Liang, Chao, et autres
Publié: (2023) -
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
par: Fan, Tiehan, et autres
Publié: (2024)