Salvato in:
| Autori principali: | Nguyen, Le Thien Phuc, Yu, Zhuoran, Lee, Yong Jae |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2501.11899 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
di: Nguyen, Le Thien Phuc, et al.
Pubblicazione: (2025)
di: Nguyen, Le Thien Phuc, et al.
Pubblicazione: (2025)
UniTalk: Towards Universal Active Speaker Detection in Real World Scenarios
di: Nguyen, Le Thien Phuc, et al.
Pubblicazione: (2025)
di: Nguyen, Le Thien Phuc, et al.
Pubblicazione: (2025)
ORFormer: Occlusion-Robust Transformer for Accurate Facial Landmark Detection
di: Chiang, Jui-Che, et al.
Pubblicazione: (2024)
di: Chiang, Jui-Che, et al.
Pubblicazione: (2024)
Evidence, Definitions and Algorithms regarding the Existence of Cohesive-Convergence Groups in Neural Network Optimization
di: Nguyen, Thien An L.
Pubblicazione: (2024)
di: Nguyen, Thien An L.
Pubblicazione: (2024)
LipSim: A Provably Robust Perceptual Similarity Metric
di: Ghazanfari, Sara, et al.
Pubblicazione: (2023)
di: Ghazanfari, Sara, et al.
Pubblicazione: (2023)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
di: Yu, Zhuoran, et al.
Pubblicazione: (2023)
di: Yu, Zhuoran, et al.
Pubblicazione: (2023)
Fiducial Focus Augmentation for Facial Landmark Detection
di: Kar, Purbayan, et al.
Pubblicazione: (2024)
di: Kar, Purbayan, et al.
Pubblicazione: (2024)
Describe Anything Model for Visual Question Answering on Text-rich Images
di: Vu, Yen-Linh, et al.
Pubblicazione: (2025)
di: Vu, Yen-Linh, et al.
Pubblicazione: (2025)
Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
di: Wu, Linzhi, et al.
Pubblicazione: (2024)
di: Wu, Linzhi, et al.
Pubblicazione: (2024)
NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition
di: Yao, Junguang, et al.
Pubblicazione: (2026)
di: Yao, Junguang, et al.
Pubblicazione: (2026)
Yo'LLaVA: Your Personalized Language and Vision Assistant
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
Landmark Stereo Dataset for Landmark Recognition and Moving Node Localization in a Non-GPS Battlefield Environment
di: Sapkota, Ganesh, et al.
Pubblicazione: (2024)
di: Sapkota, Ganesh, et al.
Pubblicazione: (2024)
Towards Universal Fake Image Detectors that Generalize Across Generative Models
di: Ojha, Utkarsh, et al.
Pubblicazione: (2023)
di: Ojha, Utkarsh, et al.
Pubblicazione: (2023)
Active Learning for Multi-class Image Classification
di: Vo, Thien Nhan
Pubblicazione: (2025)
di: Vo, Thien Nhan
Pubblicazione: (2025)
StyleLipSync: Style-based Personalized Lip-sync Video Generation
di: Ki, Taekyung, et al.
Pubblicazione: (2023)
di: Ki, Taekyung, et al.
Pubblicazione: (2023)
LipShiFT: A Certifiably Robust Shift-based Vision Transformer
di: Menon, Rohan, et al.
Pubblicazione: (2025)
di: Menon, Rohan, et al.
Pubblicazione: (2025)
Language-Assisted Feature Transformation for Anomaly Detection
di: Yun, EungGu, et al.
Pubblicazione: (2025)
di: Yun, EungGu, et al.
Pubblicazione: (2025)
Leveraging Model Soups to Classify Intangible Cultural Heritage Images from the Mekong Delta
di: Tran, Quoc-Khang, et al.
Pubblicazione: (2026)
di: Tran, Quoc-Khang, et al.
Pubblicazione: (2026)
Lightweight Facial Landmark Detection in Thermal Images via Multi-Level Cross-Modal Knowledge Transfer
di: Tong, Qiyi, et al.
Pubblicazione: (2025)
di: Tong, Qiyi, et al.
Pubblicazione: (2025)
XEdgeAI: A Human-centered Industrial Inspection Framework with Data-centric Explainable Edge AI Approach
di: Nguyen, Truong Thanh Hung, et al.
Pubblicazione: (2024)
di: Nguyen, Truong Thanh Hung, et al.
Pubblicazione: (2024)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
di: Yu, Zhuoran, et al.
Pubblicazione: (2025)
di: Yu, Zhuoran, et al.
Pubblicazione: (2025)
Boosting Facial Action Unit Detection Through Jointly Learning Facial Landmark Detection and Domain Separation and Reconstruction
di: Shang, Ziqiao, et al.
Pubblicazione: (2023)
di: Shang, Ziqiao, et al.
Pubblicazione: (2023)
End-to-end fully-binarized network design: from Generic Learned Thermometer to Block Pruning
di: Nguyen, Thien, et al.
Pubblicazione: (2025)
di: Nguyen, Thien, et al.
Pubblicazione: (2025)
A 1Mb mixed-precision quantized encoder for image classification and patch-based compression
di: Nguyen, Van Thien, et al.
Pubblicazione: (2025)
di: Nguyen, Van Thien, et al.
Pubblicazione: (2025)
MedSapiens: Taking a Pose to Rethink Medical Imaging Landmark Detection
di: Elbatel, Marawan, et al.
Pubblicazione: (2025)
di: Elbatel, Marawan, et al.
Pubblicazione: (2025)
Neonatal Face and Facial Landmark Detection from Video Recordings
di: Grooby, Ethan, et al.
Pubblicazione: (2023)
di: Grooby, Ethan, et al.
Pubblicazione: (2023)
Leveraging Expert Input for Robust and Explainable AI-Assisted Lung Cancer Detection in Chest X-rays
di: Rafferty, Amy, et al.
Pubblicazione: (2024)
di: Rafferty, Amy, et al.
Pubblicazione: (2024)
Deep Adaptation of Adult-Child Facial Expressions by Fusing Landmark Features
di: Witherow, Megan A., et al.
Pubblicazione: (2022)
di: Witherow, Megan A., et al.
Pubblicazione: (2022)
Heart Rate Classification in ECG Signals Using Machine Learning and Deep Learning
di: Vo, Thien Nhan
Pubblicazione: (2025)
di: Vo, Thien Nhan
Pubblicazione: (2025)
Geometric-Guided Few-Shot Dental Landmark Detection with Human-Centric Foundation Model
di: Wang, Anbang, et al.
Pubblicazione: (2025)
di: Wang, Anbang, et al.
Pubblicazione: (2025)
Identity-Preserving Pose-Guided Character Animation via Facial Landmarks Transformation
di: Mu, Lianrui, et al.
Pubblicazione: (2024)
di: Mu, Lianrui, et al.
Pubblicazione: (2024)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
di: Nguyen, Hieu, et al.
Pubblicazione: (2025)
di: Nguyen, Hieu, et al.
Pubblicazione: (2025)
Iterative Refinement Strategy for Automated Data Labeling: Facial Landmark Diagnosis in Medical Imaging
di: Chen, Yu-Hsi
Pubblicazione: (2024)
di: Chen, Yu-Hsi
Pubblicazione: (2024)
Follow Your Heart: Landmark-Guided Transducer Pose Scoring for Point-of-Care Echocardiography
di: Guo, Zaiyang, et al.
Pubblicazione: (2026)
di: Guo, Zaiyang, et al.
Pubblicazione: (2026)
Fast Witness Persistence for MRI Volumes via Hybrid Landmarking
di: Williams, Jorge Leonardo Ruiz
Pubblicazione: (2025)
di: Williams, Jorge Leonardo Ruiz
Pubblicazione: (2025)
LmPT: Conditional Point Transformer for Anatomical Landmark Detection on 3D Point Clouds
di: Bastico, Matteo, et al.
Pubblicazione: (2026)
di: Bastico, Matteo, et al.
Pubblicazione: (2026)
N-EIoU-YOLOv9: A Signal-Aware Bounding Box Regression Loss for Lightweight Mobile Detection of Rice Leaf Diseases
di: Duc, Dung Ta Nguyen, et al.
Pubblicazione: (2026)
di: Duc, Dung Ta Nguyen, et al.
Pubblicazione: (2026)
An Efficient and Streaming Audio Visual Active Speaker Detection System
di: Kundu, Arnav, et al.
Pubblicazione: (2024)
di: Kundu, Arnav, et al.
Pubblicazione: (2024)
Tracking-Assisted Object Detection with Event Cameras
di: Yen, Ting-Kang, et al.
Pubblicazione: (2024)
di: Yen, Ting-Kang, et al.
Pubblicazione: (2024)
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
di: Zhang, Jianrui, et al.
Pubblicazione: (2024)
di: Zhang, Jianrui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
di: Nguyen, Le Thien Phuc, et al.
Pubblicazione: (2025) -
UniTalk: Towards Universal Active Speaker Detection in Real World Scenarios
di: Nguyen, Le Thien Phuc, et al.
Pubblicazione: (2025) -
ORFormer: Occlusion-Robust Transformer for Accurate Facial Landmark Detection
di: Chiang, Jui-Che, et al.
Pubblicazione: (2024) -
Evidence, Definitions and Algorithms regarding the Existence of Cohesive-Convergence Groups in Neural Network Optimization
di: Nguyen, Thien An L.
Pubblicazione: (2024) -
LipSim: A Provably Robust Perceptual Similarity Metric
di: Ghazanfari, Sara, et al.
Pubblicazione: (2023)