Transformer with Leveraged Masked Autoencoder for video-based Pain Assessment
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Minh-Duc, Yang, Hyung-Jeong, Kim, Soo-Hyung, Shin, Ji-Eun, Kim, Seung-Won |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging WaveNet for Dynamic Listening Head Modeling from Speech
by: Nguyen, Minh-Duc, et al.
Published: (2024)
by: Nguyen, Minh-Duc, et al.
Published: (2024)
Anatomical Attention Alignment representation for Radiology Report Generation
by: Nguyen, Quang Vinh, et al.
Published: (2025)
by: Nguyen, Quang Vinh, et al.
Published: (2025)
ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
by: Vo, Hoang-Son, et al.
Published: (2025)
by: Vo, Hoang-Son, et al.
Published: (2025)
WISE-FUSE: Efficient Whole Slide Image Encoding via Coarse-to-Fine Patch Selection with VLM and LLM Knowledge Fusion
by: Shin, Yonghan, et al.
Published: (2025)
by: Shin, Yonghan, et al.
Published: (2025)
CLIMB: Controllable Longitudinal Brain Image Generation using Mamba-based Latent Diffusion Model and Gaussian-aligned Autoencoder
by: Dao, Duy-Phuong, et al.
Published: (2026)
by: Dao, Duy-Phuong, et al.
Published: (2026)
Latent Behavior Diffusion for Sequential Reaction Generation in Dyadic Setting
by: Nguyen, Minh-Duc, et al.
Published: (2025)
by: Nguyen, Minh-Duc, et al.
Published: (2025)
KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation
by: Vo-Thanh, Hoang-Son, et al.
Published: (2024)
by: Vo-Thanh, Hoang-Son, et al.
Published: (2024)
Rethinking Top Probability from Multi-view for Distracted Driver Behaviour Localization
by: Nguyen, Quang Vinh, et al.
Published: (2024)
by: Nguyen, Quang Vinh, et al.
Published: (2024)
Conditional Diffusion Model for Longitudinal Medical Image Generation
by: Dao, Duy-Phuong, et al.
Published: (2024)
by: Dao, Duy-Phuong, et al.
Published: (2024)
BoIR: Box-Supervised Instance Representation for Multi-Person Pose Estimation
by: Jeong, Uyoung, et al.
Published: (2023)
by: Jeong, Uyoung, et al.
Published: (2023)
Mask2Map: Vectorized HD Map Construction Using Bird's Eye View Segmentation Masks
by: Choi, Sehwan, et al.
Published: (2024)
by: Choi, Sehwan, et al.
Published: (2024)
SelfSwapper: Self-Supervised Face Swapping via Shape Agnostic Masked AutoEncoder
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
See It All: Contextualized Late Aggregation for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024)
by: Kim, Minjung, et al.
Published: (2024)
Self-Guided Masked Autoencoder
by: Shin, Jeongwoo, et al.
Published: (2025)
by: Shin, Jeongwoo, et al.
Published: (2025)
Cross-domain Denoising for Low-dose Multi-frame Spiral Computed Tomography
by: Lu, Yucheng, et al.
Published: (2023)
by: Lu, Yucheng, et al.
Published: (2023)
Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
by: Kim, Kinam, et al.
Published: (2025)
by: Kim, Kinam, et al.
Published: (2025)
Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial Animation
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
Leveraging Spatial Attention and Edge Context for Optimized Feature Selection in Visual Localization
by: Istighfarin, Nanda Febri, et al.
Published: (2024)
by: Istighfarin, Nanda Febri, et al.
Published: (2024)
ReffAKD: Resource-efficient Autoencoder-based Knowledge Distillation
by: Doshi, Divyang, et al.
Published: (2024)
by: Doshi, Divyang, et al.
Published: (2024)
MoST: Motion Style Transformer between Diverse Action Contents
by: Kim, Boeun, et al.
Published: (2024)
by: Kim, Boeun, et al.
Published: (2024)
Motion Cues from Image-based Point Tracking for LiDAR Scene Flow Estimation
by: Jang, Youngdong, et al.
Published: (2026)
by: Jang, Youngdong, et al.
Published: (2026)
PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation
by: Jeong, Uyoung, et al.
Published: (2025)
by: Jeong, Uyoung, et al.
Published: (2025)
ConcreTizer: Model Inversion Attack via Occupancy Classification and Dispersion Control for 3D Point Cloud Restoration
by: Kim, Youngseok, et al.
Published: (2025)
by: Kim, Youngseok, et al.
Published: (2025)
Enhanced fringe-to-phase framework using deep learning
by: Kim, Won-Hoe, et al.
Published: (2024)
by: Kim, Won-Hoe, et al.
Published: (2024)
MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
by: Oh, Youngmin, et al.
Published: (2024)
by: Oh, Youngmin, et al.
Published: (2024)
Bi-directional Contextual Attention for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024)
by: Kim, Minjung, et al.
Published: (2024)
THOM: Generating Physically Plausible Hand-Object Meshes From Text
by: Jeong, Uyoung, et al.
Published: (2026)
by: Jeong, Uyoung, et al.
Published: (2026)
Emotic Masked Autoencoder with Attention Fusion for Facial Expression Recognition
by: Nguyen-Xuan, Bach, et al.
Published: (2024)
by: Nguyen-Xuan, Bach, et al.
Published: (2024)
Improving Masked Autoencoders by Learning Where to Mask
by: Chen, Haijian, et al.
Published: (2023)
by: Chen, Haijian, et al.
Published: (2023)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
by: Jin, Hyundong, et al.
Published: (2025)
by: Jin, Hyundong, et al.
Published: (2025)
Concept Steerers: Leveraging K-Sparse Autoencoders for Test-Time Controllable Generations
by: Kim, Dahye, et al.
Published: (2025)
by: Kim, Dahye, et al.
Published: (2025)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
MINDiff: Mask-Integrated Negative Attention for Controlling Overfitting in Text-to-Image Personalization
by: Jeong, Seulgi, et al.
Published: (2025)
by: Jeong, Seulgi, et al.
Published: (2025)
Tracking the Discriminative Axis: Dual Prototypes for Test-Time OOD Detection Under Covariate Shift
by: Lee, Wooseok, et al.
Published: (2026)
by: Lee, Wooseok, et al.
Published: (2026)
3D Prior is All You Need: Cross-Task Few-shot 2D Gaze Estimation
by: Cheng, Yihua, et al.
Published: (2025)
by: Cheng, Yihua, et al.
Published: (2025)
SenseShift6D: Multimodal RGB-D Benchmarking for Robust 6D Pose Estimation across Environment and Sensor Variations
by: Han, Yegyu, et al.
Published: (2025)
by: Han, Yegyu, et al.
Published: (2025)
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
by: Hyung, Junha, et al.
Published: (2024)
by: Hyung, Junha, et al.
Published: (2024)
MAESIL: Masked Autoencoder for Enhanced Self-supervised Medical Image Learning
by: Kim, Kyeonghun, et al.
Published: (2026)
by: Kim, Kyeonghun, et al.
Published: (2026)
Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
by: Pak, Byeonghyun, et al.
Published: (2024)
by: Pak, Byeonghyun, et al.
Published: (2024)
Similar Items
-
Leveraging WaveNet for Dynamic Listening Head Modeling from Speech
by: Nguyen, Minh-Duc, et al.
Published: (2024) -
Anatomical Attention Alignment representation for Radiology Report Generation
by: Nguyen, Quang Vinh, et al.
Published: (2025) -
ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
by: Vo, Hoang-Son, et al.
Published: (2025) -
WISE-FUSE: Efficient Whole Slide Image Encoding via Coarse-to-Fine Patch Selection with VLM and LLM Knowledge Fusion
by: Shin, Yonghan, et al.
Published: (2025) -
CLIMB: Controllable Longitudinal Brain Image Generation using Mamba-based Latent Diffusion Model and Gaussian-aligned Autoencoder
by: Dao, Duy-Phuong, et al.
Published: (2026)