HI-TransPA: Hearing Impairments Translation Personal Assistant
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Zhiming, Gan, Shiyu, Zhao, Junhao, Li, Xianming, Pan, Qingyun, Wang, Peidong, Pan, Mingjun, Mo, Yuhao, Cheng, Jiajie, Chen, Chengxin, Cao, Zhonglun, Liu, Chonghan, Cheng, Shi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description
von: Pan, Mianzhi, et al.
Veröffentlicht: (2023)
von: Pan, Mianzhi, et al.
Veröffentlicht: (2023)
TeleAntiFraud-28k: An Audio-Text Slow-Thinking Dataset for Telecom Fraud Detection
von: Ma, Zhiming, et al.
Veröffentlicht: (2025)
von: Ma, Zhiming, et al.
Veröffentlicht: (2025)
Casual3DHDR: Deblurring High Dynamic Range 3D Gaussian Splatting from Casually Captured Videos
von: Gong, Shucheng, et al.
Veröffentlicht: (2025)
von: Gong, Shucheng, et al.
Veröffentlicht: (2025)
SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025)
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025)
State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026)
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026)
Beyond Isolated Utterances: Cue-Guided Interaction for Context-Dependent Conversational Multimodal Understanding
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026)
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026)
Can We Hear from Events? Generating Speech from Event Camera
von: Fang, Jingping, et al.
Veröffentlicht: (2026)
von: Fang, Jingping, et al.
Veröffentlicht: (2026)
Rethink Predicting the Optical Flow with the Kinetics Perspective
von: Cheng, Yuhao, et al.
Veröffentlicht: (2024)
von: Cheng, Yuhao, et al.
Veröffentlicht: (2024)
Rethinking Vision Transformer for Large-Scale Fine-Grained Image Retrieval
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
Interdisciplinary Translations: Sensory Perception as a Universal Language
von: Kang, Xindi, et al.
Veröffentlicht: (2024)
von: Kang, Xindi, et al.
Veröffentlicht: (2024)
MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
von: Pan, Zihan, et al.
Veröffentlicht: (2025)
von: Pan, Zihan, et al.
Veröffentlicht: (2025)
SIDQL: An Efficient Keyframe Extraction and Motion Reconstruction Framework in Motion Capture
von: Zhang, Xuling, et al.
Veröffentlicht: (2024)
von: Zhang, Xuling, et al.
Veröffentlicht: (2024)
Enhancing Few-Shot Classification without Forgetting through Multi-Level Contrastive Constraints
von: Chen, Bingzhi, et al.
Veröffentlicht: (2024)
von: Chen, Bingzhi, et al.
Veröffentlicht: (2024)
Identity-Driven Multimedia Forgery Detection via Reference Assistance
von: Xu, Junhao, et al.
Veröffentlicht: (2024)
von: Xu, Junhao, et al.
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
High Capacity Reversible Data Hiding for Encrypted 3D Mesh Models Based on Topology
von: Tang, Yun, et al.
Veröffentlicht: (2022)
von: Tang, Yun, et al.
Veröffentlicht: (2022)
Hybrid CNN-Mamba Enhancement Network for Robust Multimodal Sentiment Analysis
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
CAMERA: Adapting to Semantic Camouflage in Unsupervised Text-Attributed Graph Fraud Detection
von: Pan, Junjun, et al.
Veröffentlicht: (2026)
von: Pan, Junjun, et al.
Veröffentlicht: (2026)
Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
Towards Practical Real-Time Low-Latency Music Source Separation
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
von: Xu, Junhao, et al.
Veröffentlicht: (2025)
von: Xu, Junhao, et al.
Veröffentlicht: (2025)
RFNNS: Robust Fixed Neural Network Steganography with Universal Text-to-Image Models
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
TF-Mamba: Text-enhanced Fusion Mamba with Missing Modalities for Robust Multimodal Sentiment Analysis
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Anchoring Trends: Mitigating Social Media Popularity Prediction Drift via Feature Clustering and Expansion
von: Lee, Chia-Ming, et al.
Veröffentlicht: (2025)
von: Lee, Chia-Ming, et al.
Veröffentlicht: (2025)
Terrain Diffusion Network: Climatic-Aware Terrain Generation with Geological Sketch Guidance
von: Hu, Zexin, et al.
Veröffentlicht: (2023)
von: Hu, Zexin, et al.
Veröffentlicht: (2023)
P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark
von: Sun, Tao, et al.
Veröffentlicht: (2025)
von: Sun, Tao, et al.
Veröffentlicht: (2025)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
Joint Optimization of Buffer Delay and HARQ for Video Communications
von: Cheng, Baoping, et al.
Veröffentlicht: (2024)
von: Cheng, Baoping, et al.
Veröffentlicht: (2024)
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
von: Wang, Yuyue, et al.
Veröffentlicht: (2025)
von: Wang, Yuyue, et al.
Veröffentlicht: (2025)
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
Dual-Stream Decoupled Learning for Temporal Consistency and Speaker Interaction in AVSD
von: Xiao, Junhao, et al.
Veröffentlicht: (2025)
von: Xiao, Junhao, et al.
Veröffentlicht: (2025)
StePO-Rec: Towards Personalized Outfit Styling Assistant via Knowledge-Guided Multi-Step Reasoning
von: Bi, Yuxi, et al.
Veröffentlicht: (2025)
von: Bi, Yuxi, et al.
Veröffentlicht: (2025)
Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
EyEar: Learning Audio Synchronized Human Gaze Trajectory Based on Physics-Informed Dynamics
von: Liu, Xiaochuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaochuan, et al.
Veröffentlicht: (2025)
Mixture of Disentangled Experts with Missing Modalities for Robust Multimodal Sentiment Analysis
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
MOC-3D: Manifold-Order Consistency for Text-to-3D Generation
von: Fan, Chenyang, et al.
Veröffentlicht: (2026)
von: Fan, Chenyang, et al.
Veröffentlicht: (2026)
Visual Autoregressive Modeling for Instruction-Guided Image Editing
von: Mao, Qingyang, et al.
Veröffentlicht: (2025)
von: Mao, Qingyang, et al.
Veröffentlicht: (2025)
Stepwise Schema-Guided Prompting Framework with Parameter Efficient Instruction Tuning for Multimedia Event Extraction
von: Yuan, Xiang, et al.
Veröffentlicht: (2025)
von: Yuan, Xiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description
von: Pan, Mianzhi, et al.
Veröffentlicht: (2023) -
TeleAntiFraud-28k: An Audio-Text Slow-Thinking Dataset for Telecom Fraud Detection
von: Ma, Zhiming, et al.
Veröffentlicht: (2025) -
Casual3DHDR: Deblurring High Dynamic Range 3D Gaussian Splatting from Casually Captured Videos
von: Gong, Shucheng, et al.
Veröffentlicht: (2025) -
SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025) -
State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026)