VISTA: A Generative Egocentric Video Framework for Daily Assistance
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yu-Hsiang, Tang, Yu-Chien, Yen, An-Zi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference
by: Liou, Yow-Fu, et al.
Published: (2026)
by: Liou, Yow-Fu, et al.
Published: (2026)
Stay Hungry, Stay Foolish: On the Extended Reading Articles Generation with LLMs
by: Liou, Yow-Fu, et al.
Published: (2025)
by: Liou, Yow-Fu, et al.
Published: (2025)
MathEDU: Feedback Generation on Problem-Solving Processes for Mathematical Learning Support
by: Hsu, Wei-Ling, et al.
Published: (2025)
by: Hsu, Wei-Ling, et al.
Published: (2025)
ConceptKT: A Benchmark for Concept-Level Deficiency Prediction in Knowledge Tracing
by: Kang, Yu-Chen, et al.
Published: (2026)
by: Kang, Yu-Chen, et al.
Published: (2026)
ISSR: Iterative Selection with Self-Review for Vocabulary Test Distractor Generation
by: Liu, Yu-Cheng, et al.
Published: (2025)
by: Liu, Yu-Cheng, et al.
Published: (2025)
CARPAS: Towards Content-Aware Refinement of Provided Aspects for Summarization in Large Language Models
by: Tian, Yong-En, et al.
Published: (2025)
by: Tian, Yong-En, et al.
Published: (2025)
Fine-grained Spatiotemporal Grounding on Egocentric Videos
by: Liang, Shuo, et al.
Published: (2025)
by: Liang, Shuo, et al.
Published: (2025)
How We Refute Claims: Automatic Fact-Checking through Flaw Identification and Explanation
by: Kao, Wei-Yu, et al.
Published: (2024)
by: Kao, Wei-Yu, et al.
Published: (2024)
Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
by: Dou, Zi-Yi, et al.
Published: (2024)
by: Dou, Zi-Yi, et al.
Published: (2024)
E-QGen: Educational Lecture Abstract-based Question Generation System
by: Chen, Mao-Siang, et al.
Published: (2024)
by: Chen, Mao-Siang, et al.
Published: (2024)
A Cross-Lingual Statutory Article Retrieval Dataset for Taiwan Legal Studies
by: Wang, Yen-Hsiang, et al.
Published: (2024)
by: Wang, Yen-Hsiang, et al.
Published: (2024)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
by: Cheng, Sijie, et al.
Published: (2024)
by: Cheng, Sijie, et al.
Published: (2024)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Tree-of-Text: A Tree-based Prompting Framework for Table-to-Text Generation in the Sports Domain
by: Chiang, Shang-Hsuan, et al.
Published: (2026)
by: Chiang, Shang-Hsuan, et al.
Published: (2026)
DaMO: A Data-Efficient Multimodal Orchestrator for Temporal Reasoning with Video LLMs
by: Chiu, Bo-Cheng, et al.
Published: (2025)
by: Chiu, Bo-Cheng, et al.
Published: (2025)
MMPersuade: A Dataset and Evaluation Framework for Multimodal Persuasion
by: Qiu, Haoyi, et al.
Published: (2025)
by: Qiu, Haoyi, et al.
Published: (2025)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
by: Wang, Ziyang, et al.
Published: (2026)
by: Wang, Ziyang, et al.
Published: (2026)
Judge Like Human Examiners: A Weighted Importance Multi-Point Evaluation Framework for Generative Tasks with Long-form Answers
by: Yu, Guoxin, et al.
Published: (2026)
by: Yu, Guoxin, et al.
Published: (2026)
Automated Focused Feedback Generation for Scientific Writing Assistance
by: Chamoun, Eric, et al.
Published: (2024)
by: Chamoun, Eric, et al.
Published: (2024)
From Videos to Conversations: Egocentric Instructions for Task Assistance
by: Aggarwal, Lavisha, et al.
Published: (2026)
by: Aggarwal, Lavisha, et al.
Published: (2026)
GOSU: Retrieval-Augmented Generation with Global-Level Optimized Semantic Unit-Centric Framework
by: Zou, Xuecheng, et al.
Published: (2025)
by: Zou, Xuecheng, et al.
Published: (2025)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
by: Yoon, Jaehong, et al.
Published: (2024)
by: Yoon, Jaehong, et al.
Published: (2024)
InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
by: Yu, Chengjun, et al.
Published: (2026)
by: Yu, Chengjun, et al.
Published: (2026)
Conversational Orientation Reasoning: Egocentric-to-Allocentric Navigation with Multimodal Chain-of-Thought
by: Huang, Yu Ti
Published: (2025)
by: Huang, Yu Ti
Published: (2025)
Bailicai: A Domain-Optimized Retrieval-Augmented Generation Framework for Medical Applications
by: Long, Cui, et al.
Published: (2024)
by: Long, Cui, et al.
Published: (2024)
TranslationCorrect: A Unified Framework for Machine Translation Post-Editing with Predictive Error Assistance
by: Wasti, Syed Mekael, et al.
Published: (2025)
by: Wasti, Syed Mekael, et al.
Published: (2025)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
by: Ohkawa, Takehiko, et al.
Published: (2023)
by: Ohkawa, Takehiko, et al.
Published: (2023)
VISTA: Verification In Sequential Turn-based Assessment
by: Lewis, Ashley, et al.
Published: (2025)
by: Lewis, Ashley, et al.
Published: (2025)
LitVISTA: A Benchmark for Narrative Orchestration in Literary Text
by: Lu, Mingzhe, et al.
Published: (2026)
by: Lu, Mingzhe, et al.
Published: (2026)
The Need for a Socially-Grounded Persona Framework for User Simulation
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
by: Liu, Shaonan, et al.
Published: (2026)
by: Liu, Shaonan, et al.
Published: (2026)
Agent-Driven Large Language Models for Mandarin Lyric Generation
by: Liu, Hong-Hsiang, et al.
Published: (2024)
by: Liu, Hong-Hsiang, et al.
Published: (2024)
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
by: Suglia, Alessandro, et al.
Published: (2024)
by: Suglia, Alessandro, et al.
Published: (2024)
NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation
by: Wu, Weiming, et al.
Published: (2025)
by: Wu, Weiming, et al.
Published: (2025)
Confidence-Driven Multi-Scale Model Selection for Cost-Efficient Inference
by: Chen, Bo-Wei, et al.
Published: (2026)
by: Chen, Bo-Wei, et al.
Published: (2026)
Refining Financial Consumer Complaints through Multi-Scale Model Interaction
by: Chen, Bo-Wei, et al.
Published: (2025)
by: Chen, Bo-Wei, et al.
Published: (2025)
Paraphrase-Aligned Machine Translation
by: Chang, Ke-Ching, et al.
Published: (2024)
by: Chang, Ke-Ching, et al.
Published: (2024)
Beyond Natural Language Perplexity: Detecting Dead Code Poisoning in Code Generation Datasets
by: Tsai, Chi-Chien, et al.
Published: (2025)
by: Tsai, Chi-Chien, et al.
Published: (2025)
Similar Items
-
OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference
by: Liou, Yow-Fu, et al.
Published: (2026) -
Stay Hungry, Stay Foolish: On the Extended Reading Articles Generation with LLMs
by: Liou, Yow-Fu, et al.
Published: (2025) -
MathEDU: Feedback Generation on Problem-Solving Processes for Mathematical Learning Support
by: Hsu, Wei-Ling, et al.
Published: (2025) -
ConceptKT: A Benchmark for Concept-Level Deficiency Prediction in Knowledge Tracing
by: Kang, Yu-Chen, et al.
Published: (2026) -
ISSR: Iterative Selection with Self-Review for Vocabulary Test Distractor Generation
by: Liu, Yu-Cheng, et al.
Published: (2025)