Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Bhattacharya, Uttaran, Wu, Gang, Petrangeli, Stefano, Swaminathan, Viswanathan, Manocha, Dinesh |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HighlightMe: Detecting Highlights from Human-Centric Videos
by: Bhattacharya, Uttaran, et al.
Published: (2021)
by: Bhattacharya, Uttaran, et al.
Published: (2021)
Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
by: Bhattacharya, Uttaran, et al.
Published: (2024)
by: Bhattacharya, Uttaran, et al.
Published: (2024)
Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning
by: Bhattacharya, Uttaran, et al.
Published: (2021)
by: Bhattacharya, Uttaran, et al.
Published: (2021)
Proc3D: Procedural 3D Generation and Parametric Editing of 3D Shapes with Large Language Models
by: Raji, Fadlullah, et al.
Published: (2026)
by: Raji, Fadlullah, et al.
Published: (2026)
Take an Emotion Walk: Perceiving Emotions from Gaits Using Hierarchical Attention Pooling and Affective Mapping
by: Bhattacharya, Uttaran, et al.
Published: (2019)
by: Bhattacharya, Uttaran, et al.
Published: (2019)
Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents
by: Bhattacharya, Uttaran, et al.
Published: (2021)
by: Bhattacharya, Uttaran, et al.
Published: (2021)
Generating Emotive Gaits for Virtual Agents Using Affect-Based Autoregression
by: Bhattacharya, Uttaran, et al.
Published: (2020)
by: Bhattacharya, Uttaran, et al.
Published: (2020)
STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits
by: Bhattacharya, Uttaran, et al.
Published: (2019)
by: Bhattacharya, Uttaran, et al.
Published: (2019)
Large Content And Behavior Models To Understand, Simulate, And Optimize Content And Behavior
by: Khandelwal, Ashmit, et al.
Published: (2023)
by: Khandelwal, Ashmit, et al.
Published: (2023)
Inst4DGS: Instance-Decomposed 4D Gaussian Splatting with Multi-Video Label Permutation Learning
by: Lee, Yonghan, et al.
Published: (2026)
by: Lee, Yonghan, et al.
Published: (2026)
Show Me How: Benefits and Challenges of Agent-Augmented Counterfactual Explanations for Non-Expert Users
by: Bhattacharya, Aditya, et al.
Published: (2025)
by: Bhattacharya, Aditya, et al.
Published: (2025)
"May I Speak?": Multi-modal Attention Guidance in Social VR Group Conversations
by: Lee, Geonsun, et al.
Published: (2024)
by: Lee, Geonsun, et al.
Published: (2024)
DMCA: Dense Multi-agent Navigation using Attention and Communication
by: Arul, Senthil Hariharan, et al.
Published: (2022)
by: Arul, Senthil Hariharan, et al.
Published: (2022)
From Pixels to Policies: Reinforcing Spatial Reasoning in Language Models for Content-Aware Layout Design
by: Li, Sha, et al.
Published: (2026)
by: Li, Sha, et al.
Published: (2026)
ScreenLLM: Stateful Screen Schema for Efficient Action Understanding and Prediction
by: Jin, Yiqiao, et al.
Published: (2025)
by: Jin, Yiqiao, et al.
Published: (2025)
EM-GANSim: Real-time and Accurate EM Simulation Using Conditional GANs for 3D Indoor Scenes
by: Wang, Ruichen, et al.
Published: (2024)
by: Wang, Ruichen, et al.
Published: (2024)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
by: Cheng, Stephen, et al.
Published: (2026)
by: Cheng, Stephen, et al.
Published: (2026)
Differentiable Frequency-based Disentanglement for Aerial Video Action Recognition
by: Kothandaraman, Divya, et al.
Published: (2022)
by: Kothandaraman, Divya, et al.
Published: (2022)
Can LLMs Generate Human-Like Wayfinding Instructions? Towards Platform-Agnostic Embodied Instruction Synthesis
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
MemCtrl: Using MLLMs as Active Memory Controllers on Embodied Agents
by: Dorbala, Vishnu Sashank, et al.
Published: (2026)
by: Dorbala, Vishnu Sashank, et al.
Published: (2026)
LLMs as Layout Designers: Enhanced Spatial Reasoning for Content-Aware Layout Generation
by: Li, Sha, et al.
Published: (2025)
by: Li, Sha, et al.
Published: (2025)
Efficient and Robust Registration on the 3D Special Euclidean Group
by: Bhattacharya, Uttaran, et al.
Published: (2019)
by: Bhattacharya, Uttaran, et al.
Published: (2019)
GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling
by: Shekhar, Shivanshu, et al.
Published: (2026)
by: Shekhar, Shivanshu, et al.
Published: (2026)
Token-Level Adversarial Prompt Detection Based on Perplexity Measures and Contextual Information
by: Hu, Zhengmian, et al.
Published: (2023)
by: Hu, Zhengmian, et al.
Published: (2023)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
by: Anand, Nishit, et al.
Published: (2024)
by: Anand, Nishit, et al.
Published: (2024)
LoLep: Single-View View Synthesis with Locally-Learned Planes and Self-Attention Occlusion Inference
by: Wang, Cong, et al.
Published: (2023)
by: Wang, Cong, et al.
Published: (2023)
DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
by: Bhattacharya, Aneesh, et al.
Published: (2023)
by: Bhattacharya, Aneesh, et al.
Published: (2023)
Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces
by: Chowdhury, Sanjoy, et al.
Published: (2026)
by: Chowdhury, Sanjoy, et al.
Published: (2026)
PACE: Data-Driven Virtual Agent Interaction in Dense and Cluttered Environments
by: Mullen, James, et al.
Published: (2023)
by: Mullen, James, et al.
Published: (2023)
Listen2Scene: Interactive material-aware binaural sound propagation for reconstructed 3D scenes
by: Ratnarajah, Anton, et al.
Published: (2023)
by: Ratnarajah, Anton, et al.
Published: (2023)
A Multi-Head Attention Soft Random Forest for Interpretable Patient No-Show Prediction
by: Amalina, Ninda Nurseha, et al.
Published: (2025)
by: Amalina, Ninda Nurseha, et al.
Published: (2025)
Show Me What I Should Know! Active, Contextual Learning on the Job--A Review Essay.
by: Landgren, Craig Randall
Published: (1993)
by: Landgren, Craig Randall
Published: (1993)
Text Prompting for Multi-Concept Video Customization by Autoregressive Generation
by: Kothandaraman, Divya, et al.
Published: (2024)
by: Kothandaraman, Divya, et al.
Published: (2024)
Show Me What and Where has Changed? Question Answering and Grounding for Remote Sensing Change Detection
by: Li, Ke, et al.
Published: (2024)
by: Li, Ke, et al.
Published: (2024)
MeLFusion: Synthesizing Music from Image and Language Cues using Diffusion Models
by: Chowdhury, Sanjoy, et al.
Published: (2024)
by: Chowdhury, Sanjoy, et al.
Published: (2024)
Towards Optimal Multi-draft Speculative Decoding
by: Hu, Zhengmian, et al.
Published: (2025)
by: Hu, Zhengmian, et al.
Published: (2025)
UnMA-CapSumT: Unified and Multi-Head Attention-driven Caption Summarization Transformer
by: Sharma, Dhruv, et al.
Published: (2024)
by: Sharma, Dhruv, et al.
Published: (2024)
SyncTrack4D: Cross-Video Motion Alignment and Video Synchronization for Multi-Video 4D Gaussian Splatting
by: Lee, Yonghan, et al.
Published: (2025)
by: Lee, Yonghan, et al.
Published: (2025)
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
by: Pu, Yujiang, et al.
Published: (2025)
by: Pu, Yujiang, et al.
Published: (2025)
"Show Me What's Wrong!": Combining Charts and Text to Guide Data Analysis
by: Feliciano, Beatriz, et al.
Published: (2024)
by: Feliciano, Beatriz, et al.
Published: (2024)
Similar Items
-
HighlightMe: Detecting Highlights from Human-Centric Videos
by: Bhattacharya, Uttaran, et al.
Published: (2021) -
Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
by: Bhattacharya, Uttaran, et al.
Published: (2024) -
Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning
by: Bhattacharya, Uttaran, et al.
Published: (2021) -
Proc3D: Procedural 3D Generation and Parametric Editing of 3D Shapes with Large Language Models
by: Raji, Fadlullah, et al.
Published: (2026) -
Take an Emotion Walk: Perceiving Emotions from Gaits Using Hierarchical Attention Pooling and Affective Mapping
by: Bhattacharya, Uttaran, et al.
Published: (2019)