SwissADT: An Audio Description Translation System for Swiss Languages
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fischer, Lukas, Gao, Yingqiang, Lintner, Alexa, Ebling, Sarah |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies
von: Gao, Yingqiang, et al.
Veröffentlicht: (2024)
von: Gao, Yingqiang, et al.
Veröffentlicht: (2024)
Benchmarking NLP-supported Language Sample Analysis for Swiss Children's Speech
von: Ryser, Anja, et al.
Veröffentlicht: (2025)
von: Ryser, Anja, et al.
Veröffentlicht: (2025)
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
von: Khalil, Mohammad Amer, et al.
Veröffentlicht: (2026)
von: Khalil, Mohammad Amer, et al.
Veröffentlicht: (2026)
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
von: Tsangko, Iosif, et al.
Veröffentlicht: (2026)
von: Tsangko, Iosif, et al.
Veröffentlicht: (2026)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
ASI-Seg: Audio-Driven Surgical Instrument Segmentation with Surgeon Intention Understanding
von: Chen, Zhen, et al.
Veröffentlicht: (2024)
von: Chen, Zhen, et al.
Veröffentlicht: (2024)
Investigating Disability Representations in Text-to-Image Models
von: Tian, Yang, et al.
Veröffentlicht: (2026)
von: Tian, Yang, et al.
Veröffentlicht: (2026)
Can Large Language Models Capture Video Game Engagement?
von: Melhart, David, et al.
Veröffentlicht: (2025)
von: Melhart, David, et al.
Veröffentlicht: (2025)
GUICourse: From General Vision Language Models to Versatile GUI Agents
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
Forest-Chat: Adapting Vision-Language Agents for Interactive Forest Change Analysis
von: Brock, James, et al.
Veröffentlicht: (2026)
von: Brock, James, et al.
Veröffentlicht: (2026)
Seeing Eye to AI: Human Alignment via Gaze-Based Response Rewards for Large Language Models
von: Lopez-Cardona, Angela, et al.
Veröffentlicht: (2024)
von: Lopez-Cardona, Angela, et al.
Veröffentlicht: (2024)
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
von: Lin, Ronghao, et al.
Veröffentlicht: (2025)
von: Lin, Ronghao, et al.
Veröffentlicht: (2025)
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
von: Mukhopadhyay, Srija, et al.
Veröffentlicht: (2024)
von: Mukhopadhyay, Srija, et al.
Veröffentlicht: (2024)
Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
von: Huang, Ting-Hao 'Kenneth', et al.
Veröffentlicht: (2025)
von: Huang, Ting-Hao 'Kenneth', et al.
Veröffentlicht: (2025)
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
von: Xie, Tianbao, et al.
Veröffentlicht: (2025)
von: Xie, Tianbao, et al.
Veröffentlicht: (2025)
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
DesignPref: Capturing Personal Preferences in Visual Design Generation
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2025)
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2025)
AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent
von: Fan, Jingru, et al.
Veröffentlicht: (2025)
von: Fan, Jingru, et al.
Veröffentlicht: (2025)
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2026)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2026)
Measuring Agreeableness Bias in Multimodal Models
von: Lim, Jaehyuk, et al.
Veröffentlicht: (2024)
von: Lim, Jaehyuk, et al.
Veröffentlicht: (2024)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
Tur[k]ingBench: A Challenge Benchmark for Web Agents
von: Xu, Kevin, et al.
Veröffentlicht: (2024)
von: Xu, Kevin, et al.
Veröffentlicht: (2024)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
History-Aware Reasoning for GUI Agents
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
Code2World: A GUI World Model via Renderable Code Generation
von: Zheng, Yuhao, et al.
Veröffentlicht: (2026)
von: Zheng, Yuhao, et al.
Veröffentlicht: (2026)
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
von: Chen, Xuetian, et al.
Veröffentlicht: (2025)
von: Chen, Xuetian, et al.
Veröffentlicht: (2025)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
von: Kapoor, Raghav, et al.
Veröffentlicht: (2024)
von: Kapoor, Raghav, et al.
Veröffentlicht: (2024)
A Survey on (M)LLM-Based GUI Agents
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild
von: Chen, Baiyu, et al.
Veröffentlicht: (2026)
von: Chen, Baiyu, et al.
Veröffentlicht: (2026)
ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles
von: Yin, Kayo, et al.
Veröffentlicht: (2024)
von: Yin, Kayo, et al.
Veröffentlicht: (2024)
How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?
von: Lo, Leo Yu-Ho, et al.
Veröffentlicht: (2024)
von: Lo, Leo Yu-Ho, et al.
Veröffentlicht: (2024)
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies
von: Gao, Yingqiang, et al.
Veröffentlicht: (2024) -
Benchmarking NLP-supported Language Sample Analysis for Swiss Children's Speech
von: Ryser, Anja, et al.
Veröffentlicht: (2025) -
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
von: Khalil, Mohammad Amer, et al.
Veröffentlicht: (2026) -
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
von: Tsangko, Iosif, et al.
Veröffentlicht: (2026) -
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)