Guardado en:
| Autores principales: | Liu, Qingyuan, Tsai, Yun-Yun, Zha, Ruijian, Li, Victoria, Shi, Pengyuan, Mao, Chengzhi, Yang, Junfeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2502.14994 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Turns Out I'm Not Real: Towards Robust Detection of AI-Generated Videos
por: Liu, Qingyuan, et al.
Publicado: (2024)
por: Liu, Qingyuan, et al.
Publicado: (2024)
GDA: Generalized Diffusion for Robust Test-time Adaptation
por: Tsai, Yun-Yun, et al.
Publicado: (2024)
por: Tsai, Yun-Yun, et al.
Publicado: (2024)
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
por: Song, Python, et al.
Publicado: (2025)
por: Song, Python, et al.
Publicado: (2025)
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
por: Huang, Yiyang, et al.
Publicado: (2025)
por: Huang, Yiyang, et al.
Publicado: (2025)
AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
por: Ahn, Sunghyun, et al.
Publicado: (2025)
por: Ahn, Sunghyun, et al.
Publicado: (2025)
Robust and Annotation-Free Wound Segmentation on Noisy Real-World Pressure Ulcer Images: Towards Automated DESIGN-R\textsuperscript{\textregistered} Assessment
por: Tsai, Yun-Cheng
Publicado: (2025)
por: Tsai, Yun-Cheng
Publicado: (2025)
ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models
por: Fang, Zixun, et al.
Publicado: (2025)
por: Fang, Zixun, et al.
Publicado: (2025)
VanGogh: A Unified Multimodal Diffusion-based Framework for Video Colorization
por: Fang, Zixun, et al.
Publicado: (2025)
por: Fang, Zixun, et al.
Publicado: (2025)
Detecting AI-Generated Video via Frame Consistency
por: Ma, Long, et al.
Publicado: (2024)
por: Ma, Long, et al.
Publicado: (2024)
MedM-VL: What Makes a Good Medical LVLM?
por: Shi, Yiming, et al.
Publicado: (2025)
por: Shi, Yiming, et al.
Publicado: (2025)
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
por: Peng, Liyang, et al.
Publicado: (2025)
por: Peng, Liyang, et al.
Publicado: (2025)
LVLM-Composer's Explicit Planning for Image Generation
por: Ramsey, Spencer, et al.
Publicado: (2025)
por: Ramsey, Spencer, et al.
Publicado: (2025)
DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion
por: Long, Yongji, et al.
Publicado: (2026)
por: Long, Yongji, et al.
Publicado: (2026)
EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video Reconstruction
por: Ge, Chengjie, et al.
Publicado: (2025)
por: Ge, Chengjie, et al.
Publicado: (2025)
Generative Neural Video Compression via Video Diffusion Prior
por: Mao, Qi, et al.
Publicado: (2025)
por: Mao, Qi, et al.
Publicado: (2025)
Constructing a 3D Scene from a Single Image
por: Zheng, Kaizhi, et al.
Publicado: (2025)
por: Zheng, Kaizhi, et al.
Publicado: (2025)
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
por: Zhang, Chenshuang, et al.
Publicado: (2025)
por: Zhang, Chenshuang, et al.
Publicado: (2025)
ViViD: Video Virtual Try-on using Diffusion Models
por: Fang, Zixun, et al.
Publicado: (2024)
por: Fang, Zixun, et al.
Publicado: (2024)
Fighting Hallucinations with Counterfactuals: Diffusion-Guided Perturbations for LVLM Hallucination Suppression
por: Dastmalchi, Hamidreza, et al.
Publicado: (2026)
por: Dastmalchi, Hamidreza, et al.
Publicado: (2026)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
por: Hwang, Yerin, et al.
Publicado: (2025)
por: Hwang, Yerin, et al.
Publicado: (2025)
Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks
por: Kwon, JuneHyoung, et al.
Publicado: (2026)
por: Kwon, JuneHyoung, et al.
Publicado: (2026)
AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLM
por: Zhong, Li'an, et al.
Publicado: (2026)
por: Zhong, Li'an, et al.
Publicado: (2026)
An Efficient Streaming Video Understanding Framework with Agentic Control
por: Liu, Jinming, et al.
Publicado: (2026)
por: Liu, Jinming, et al.
Publicado: (2026)
Mitty: Diffusion-based Human-to-Robot Video Generation
por: Song, Yiren, et al.
Publicado: (2025)
por: Song, Yiren, et al.
Publicado: (2025)
Towards Better De-raining Generalization via Rainy Characteristics Memorization and Replay
por: Wang, Kunyu, et al.
Publicado: (2025)
por: Wang, Kunyu, et al.
Publicado: (2025)
ReManNet: A Riemannian Manifold Network for Monocular 3D Lane Detection
por: Hong, Chengzhi, et al.
Publicado: (2026)
por: Hong, Chengzhi, et al.
Publicado: (2026)
SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation
por: He, Yun, et al.
Publicado: (2026)
por: He, Yun, et al.
Publicado: (2026)
VALD: Multi-Stage Vision Attack Detection for Efficient LVLM Defense
por: Kadvil, Nadav, et al.
Publicado: (2026)
por: Kadvil, Nadav, et al.
Publicado: (2026)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
por: Jin, Hongbo, et al.
Publicado: (2025)
por: Jin, Hongbo, et al.
Publicado: (2025)
Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning
por: Wang, Yifan, et al.
Publicado: (2026)
por: Wang, Yifan, et al.
Publicado: (2026)
ExpertGen: Training-Free Expert Guidance for Controllable Text-to-Face Generation
por: Shi, Liang, et al.
Publicado: (2025)
por: Shi, Liang, et al.
Publicado: (2025)
VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
por: Yang, Lehan, et al.
Publicado: (2025)
por: Yang, Lehan, et al.
Publicado: (2025)
HERO: Human Reaction Generation from Videos
por: Yu, Chengjun, et al.
Publicado: (2025)
por: Yu, Chengjun, et al.
Publicado: (2025)
NEWTON: Agentic Planning for Physically Grounded Video Generation
por: Feng, Yuxiang, et al.
Publicado: (2026)
por: Feng, Yuxiang, et al.
Publicado: (2026)
Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
por: Wu, Yuanchen, et al.
Publicado: (2025)
por: Wu, Yuanchen, et al.
Publicado: (2025)
Efficient Test-time Adaptive Object Detection via Sensitivity-Guided Pruning
por: Wang, Kunyu, et al.
Publicado: (2025)
por: Wang, Kunyu, et al.
Publicado: (2025)
Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
por: Li, Hang, et al.
Publicado: (2023)
por: Li, Hang, et al.
Publicado: (2023)
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
por: Li, Jie, et al.
Publicado: (2023)
por: Li, Jie, et al.
Publicado: (2023)
Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing
por: Zhang, Xu, et al.
Publicado: (2025)
por: Zhang, Xu, et al.
Publicado: (2025)
Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation
por: Garcia, Fernando Gabriela, et al.
Publicado: (2025)
por: Garcia, Fernando Gabriela, et al.
Publicado: (2025)
Ejemplares similares
-
Turns Out I'm Not Real: Towards Robust Detection of AI-Generated Videos
por: Liu, Qingyuan, et al.
Publicado: (2024) -
GDA: Generalized Diffusion for Robust Test-time Adaptation
por: Tsai, Yun-Yun, et al.
Publicado: (2024) -
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
por: Song, Python, et al.
Publicado: (2025) -
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
por: Huang, Yiyang, et al.
Publicado: (2025) -
AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
por: Ahn, Sunghyun, et al.
Publicado: (2025)