HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Faure, Gueter Josmy, Yeh, Jia-Fong, Chen, Min-Hung, Su, Hung-Ting, Lai, Shang-Hong, Hsu, Winston H. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
von: Chunhachatrachai, Pawat, et al.
Veröffentlicht: (2026)
von: Chunhachatrachai, Pawat, et al.
Veröffentlicht: (2026)
MovieCORE: COgnitive REasoning in Movies
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2025)
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2025)
SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization
von: Chen, Posheng, et al.
Veröffentlicht: (2026)
von: Chen, Posheng, et al.
Veröffentlicht: (2026)
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
von: Chen, Pei-An, et al.
Veröffentlicht: (2026)
von: Chen, Pei-An, et al.
Veröffentlicht: (2026)
Tracking-Assisted Object Detection with Event Cameras
von: Yen, Ting-Kang, et al.
Veröffentlicht: (2024)
von: Yen, Ting-Kang, et al.
Veröffentlicht: (2024)
Revisiting Semi-supervised Adversarial Robustness via Noise-aware Online Robust Distillation
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
Context-Aware Replanning with Pre-explored Semantic Map for Object Navigation
von: Ko, Po-Chen, et al.
Veröffentlicht: (2024)
von: Ko, Po-Chen, et al.
Veröffentlicht: (2024)
Investigating Video Reasoning Capability of Large Language Models with Tropes in Movies
von: Su, Hung-Ting, et al.
Veröffentlicht: (2024)
von: Su, Hung-Ting, et al.
Veröffentlicht: (2024)
Tel2Veh: Fusion of Telecom Data and Vehicle Flow to Predict Camera-Free Traffic via a Spatio-Temporal Framework
von: Lin, ChungYi, et al.
Veröffentlicht: (2024)
von: Lin, ChungYi, et al.
Veröffentlicht: (2024)
Distribution Discrepancy and Feature Heterogeneity for Active 3D Object Detection
von: Chen, Huang-Yu, et al.
Veröffentlicht: (2024)
von: Chen, Huang-Yu, et al.
Veröffentlicht: (2024)
Unveiling Narrative Reasoning Limits of Large Language Models with Trope in Movie Synopses
von: Su, Hung-Ting, et al.
Veröffentlicht: (2024)
von: Su, Hung-Ting, et al.
Veröffentlicht: (2024)
Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation
von: Lin, Tzu-Jung, et al.
Veröffentlicht: (2025)
von: Lin, Tzu-Jung, et al.
Veröffentlicht: (2025)
Large Language Models Prompting With Episodic Memory
von: Do, Dai, et al.
Veröffentlicht: (2024)
von: Do, Dai, et al.
Veröffentlicht: (2024)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
von: Huang, Wei-Jhe, et al.
Veröffentlicht: (2024)
von: Huang, Wei-Jhe, et al.
Veröffentlicht: (2024)
A$^2$TG: Adaptive Anisotropic Textured Gaussians for Efficient 3D Scene Representation
von: Hsu, Sheng-Chi, et al.
Veröffentlicht: (2026)
von: Hsu, Sheng-Chi, et al.
Veröffentlicht: (2026)
VICtoR: Learning Hierarchical Vision-Instruction Correlation Rewards for Long-horizon Manipulation
von: Hung, Kuo-Han, et al.
Veröffentlicht: (2024)
von: Hung, Kuo-Han, et al.
Veröffentlicht: (2024)
AED: Adaptable Error Detection for Few-shot Imitation Policy
von: Yeh, Jia-Fong, et al.
Veröffentlicht: (2024)
von: Yeh, Jia-Fong, et al.
Veröffentlicht: (2024)
From Persona to Person: Enhancing the Naturalness with Multiple Discourse Relations Graph Learning in Personalized Dialogue Generation
von: Hsu, Chih-Hao, et al.
Veröffentlicht: (2025)
von: Hsu, Chih-Hao, et al.
Veröffentlicht: (2025)
Semantic Prompt Learning for Weakly-Supervised Semantic Segmentation
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2024)
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2024)
Contrastive Bi-Projector for Unsupervised Domain Adaption
von: Huang, Lin-Chieh, et al.
Veröffentlicht: (2023)
von: Huang, Lin-Chieh, et al.
Veröffentlicht: (2023)
Shared-unique Features and Task-aware Prioritized Sampling on Multi-task Reinforcement Learning
von: Lin, Po-Shao, et al.
Veröffentlicht: (2024)
von: Lin, Po-Shao, et al.
Veröffentlicht: (2024)
VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
von: Cheng, Ying, et al.
Veröffentlicht: (2025)
von: Cheng, Ying, et al.
Veröffentlicht: (2025)
XAI-Enhanced Semantic Segmentation Models for Visual Quality Inspection
von: Clement, Tobias, et al.
Veröffentlicht: (2024)
von: Clement, Tobias, et al.
Veröffentlicht: (2024)
ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
CTGAN: Semantic-guided Conditional Texture Generator for 3D Shapes
von: Pan, Yi-Ting, et al.
Veröffentlicht: (2024)
von: Pan, Yi-Ting, et al.
Veröffentlicht: (2024)
Open-World Evaluation for Retrieving Diverse Perspectives
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2024)
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2024)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
von: Ting, Hsu-Tzu, et al.
Veröffentlicht: (2025)
von: Ting, Hsu-Tzu, et al.
Veröffentlicht: (2025)
Can LLMs Understand the Implication of Emphasized Sentences in Dialogue?
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
GSQA: An End-to-End Model for Generative Spoken Question Answering
von: Shih, Min-Han, et al.
Veröffentlicht: (2023)
von: Shih, Min-Han, et al.
Veröffentlicht: (2023)
Unsupervised Image Prior via Prompt Learning and CLIP Semantic Guidance for Low-Light Image Enhancement
von: Morawski, Igor, et al.
Veröffentlicht: (2024)
von: Morawski, Igor, et al.
Veröffentlicht: (2024)
Exploring Scholarly Data by Semantic Query on Knowledge Graph Embedding Space
von: Tran, Hung Nghiep, et al.
Veröffentlicht: (2019)
von: Tran, Hung Nghiep, et al.
Veröffentlicht: (2019)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
von: Zhang, Haowei, et al.
Veröffentlicht: (2026)
von: Zhang, Haowei, et al.
Veröffentlicht: (2026)
RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
von: Hsu, Chia-Hsuan, et al.
Veröffentlicht: (2025)
von: Hsu, Chia-Hsuan, et al.
Veröffentlicht: (2025)
Geographic Blind Spots in AI Control Monitors: A Cross-National Audit of Claude Opus 4.6
von: Hung, Jason
Veröffentlicht: (2026)
von: Hung, Jason
Veröffentlicht: (2026)
CFEVER: A Chinese Fact Extraction and VERification Dataset
von: Lin, Ying-Jia, et al.
Veröffentlicht: (2024)
von: Lin, Ying-Jia, et al.
Veröffentlicht: (2024)
Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2023)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2023)
LoBE-GS: Load-Balanced and Efficient 3D Gaussian Splatting for Large-Scale Scene Reconstruction
von: Hung, Sheng-Hsiang, et al.
Veröffentlicht: (2025)
von: Hung, Sheng-Hsiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026) -
SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
von: Chunhachatrachai, Pawat, et al.
Veröffentlicht: (2026) -
MovieCORE: COgnitive REasoning in Movies
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2025) -
SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization
von: Chen, Posheng, et al.
Veröffentlicht: (2026) -
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)