VISTA: Verification In Sequential Turn-based Assessment
Fuente:
arXiv
Saved in:
| Main Authors: | Lewis, Ashley, Perrault, Andrew, Fosler-Lussier, Eric, White, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances
by: Chun, Jiyun, et al.
Published: (2026)
by: Chun, Jiyun, et al.
Published: (2026)
Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers
by: Serai, Prashant, et al.
Published: (2024)
by: Serai, Prashant, et al.
Published: (2024)
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
by: Jones, Jaylen, et al.
Published: (2024)
by: Jones, Jaylen, et al.
Published: (2024)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
by: Ginjala, Srishti, et al.
Published: (2026)
by: Ginjala, Srishti, et al.
Published: (2026)
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
by: Liao, Zeyi, et al.
Published: (2025)
by: Liao, Zeyi, et al.
Published: (2025)
Beyond the Black Box: Demystifying Multi-Turn LLM Reasoning with VISTA
by: Zhang, Yiran, et al.
Published: (2025)
by: Zhang, Yiran, et al.
Published: (2025)
Music on the Move
by: Fosler-Lussier, Danielle
Published: (2020)
by: Fosler-Lussier, Danielle
Published: (2020)
Improving Transducer-Based Spoken Language Understanding with Self-Conditioned CTC and Knowledge Transfer
by: Sunder, Vishal, et al.
Published: (2025)
by: Sunder, Vishal, et al.
Published: (2025)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
by: Jones, Jaylen, et al.
Published: (2026)
by: Jones, Jaylen, et al.
Published: (2026)
NaturalTurn: A Method to Segment Speech into Psychologically Meaningful Conversational Turns
by: Cooney, Gus, et al.
Published: (2024)
by: Cooney, Gus, et al.
Published: (2024)
End-to-End Diarization utilizing Attractor Deep Clustering
by: Palzer, David, et al.
Published: (2025)
by: Palzer, David, et al.
Published: (2025)
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
by: Palzer, David, et al.
Published: (2025)
by: Palzer, David, et al.
Published: (2025)
Assessing the Reasoning Capabilities of LLMs in the context of Evidence-based Claim Verification
by: Dougrez-Lewis, John, et al.
Published: (2024)
by: Dougrez-Lewis, John, et al.
Published: (2024)
Towards a Perspectivist Turn in Argument Quality Assessment
by: Romberg, Julia, et al.
Published: (2025)
by: Romberg, Julia, et al.
Published: (2025)
ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models
by: Owiredu-Ashley, Harry
Published: (2026)
by: Owiredu-Ashley, Harry
Published: (2026)
From Myopic Selection to Long-Horizon Awareness: Sequential LLM Routing for Multi-Turn Dialogue
by: Zhang, Jiarui, et al.
Published: (2026)
by: Zhang, Jiarui, et al.
Published: (2026)
Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents
by: Lewis, Ashley, et al.
Published: (2025)
by: Lewis, Ashley, et al.
Published: (2025)
More Rounds, More Noise: Why Multi-Turn Review Fails to Improve Cross-Context Verification
by: Tae-Eun, Song
Published: (2026)
by: Tae-Eun, Song
Published: (2026)
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback
by: Byun, Ju-Seung, et al.
Published: (2024)
by: Byun, Ju-Seung, et al.
Published: (2024)
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
by: Chen, Jingyi, et al.
Published: (2025)
by: Chen, Jingyi, et al.
Published: (2025)
VISTA: A Generative Egocentric Video Framework for Daily Assistance
by: Liu, Yu-Hsiang, et al.
Published: (2026)
by: Liu, Yu-Hsiang, et al.
Published: (2026)
Auditing Support Strategies in LLMs through Grounded Multi-Turn Social Simulation
by: Star, Michelle, et al.
Published: (2026)
by: Star, Michelle, et al.
Published: (2026)
A Browser-based Open Source Assistant for Multimodal Content Verification
by: Milner, Rosanna, et al.
Published: (2026)
by: Milner, Rosanna, et al.
Published: (2026)
Evaluating Step-by-Step Reasoning through Symbolic Verification
by: Zhang, Yi-Fan, et al.
Published: (2022)
by: Zhang, Yi-Fan, et al.
Published: (2022)
Eliciting Behaviors in Multi-Turn Conversations
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
A Graph-based Verification Framework for Fact-Checking
by: Huang, Yani, et al.
Published: (2025)
by: Huang, Yani, et al.
Published: (2025)
ChronoFact: Timeline-based Temporal Fact Verification
by: Barik, Anab Maulana, et al.
Published: (2024)
by: Barik, Anab Maulana, et al.
Published: (2024)
Self-Verification is All You Need To Pass The Japanese Bar Examination
by: Shin, Andrew
Published: (2026)
by: Shin, Andrew
Published: (2026)
MUSIC: MUlti-Step Instruction Contrast for Multi-Turn Reward Models
by: Li, Wenzhe, et al.
Published: (2025)
by: Li, Wenzhe, et al.
Published: (2025)
VISTA: Visualization of Token Attribution via Efficient Analysis
by: Ahmed, Syed, et al.
Published: (2026)
by: Ahmed, Syed, et al.
Published: (2026)
Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs
by: Nguyen, Minh Nhat, et al.
Published: (2024)
by: Nguyen, Minh Nhat, et al.
Published: (2024)
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM
by: Kiruluta, Andrew, et al.
Published: (2025)
by: Kiruluta, Andrew, et al.
Published: (2025)
Sequential Diagnosis with Language Models
by: Nori, Harsha, et al.
Published: (2025)
by: Nori, Harsha, et al.
Published: (2025)
Modeling Future Conversation Turns to Teach LLMs to Ask Clarifying Questions
by: Zhang, Michael J. Q., et al.
Published: (2024)
by: Zhang, Michael J. Q., et al.
Published: (2024)
A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping
by: Chen, Dingwei, et al.
Published: (2026)
by: Chen, Dingwei, et al.
Published: (2026)
EICAP: Deep Dive in Assessment and Enhancement of Large Language Models in Emotional Intelligence through Multi-Turn Conversations
by: Nazar, Nizi, et al.
Published: (2025)
by: Nazar, Nizi, et al.
Published: (2025)
LitVISTA: A Benchmark for Narrative Orchestration in Literary Text
by: Lu, Mingzhe, et al.
Published: (2026)
by: Lu, Mingzhe, et al.
Published: (2026)
Multi-domain Multilingual Sentiment Analysis in Industry: Predicting Aspect-based Opinion Quadruples
by: White, Benjamin, et al.
Published: (2025)
by: White, Benjamin, et al.
Published: (2025)
LLMForecaster: Improving Seasonal Event Forecasts with Unstructured Textual Data
by: Zhang, Hanyu, et al.
Published: (2024)
by: Zhang, Hanyu, et al.
Published: (2024)
Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions
by: Kim, JiWoo, et al.
Published: (2025)
by: Kim, JiWoo, et al.
Published: (2025)
Similar Items
-
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances
by: Chun, Jiyun, et al.
Published: (2026) -
Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers
by: Serai, Prashant, et al.
Published: (2024) -
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
by: Jones, Jaylen, et al.
Published: (2024) -
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
by: Ginjala, Srishti, et al.
Published: (2026) -
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
by: Liao, Zeyi, et al.
Published: (2025)