CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Guangzhi, Manakul, Potsawee, Liusie, Adian, Pipatanakul, Kunat, Zhang, Chao, Woodland, Phil, Gales, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
by: Liusie, Adian, et al.
Published: (2023)
by: Liusie, Adian, et al.
Published: (2023)
SkillAggregation: Reference-free LLM-Dependent Aggregation
by: Sun, Guangzhi, et al.
Published: (2024)
by: Sun, Guangzhi, et al.
Published: (2024)
Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models
by: Manakul, Potsawee, et al.
Published: (2024)
by: Manakul, Potsawee, et al.
Published: (2024)
Adapting Language-Specific LLMs to a Reasoning Model in One Day via Model Merging -- An Open Recipe
by: Pipatanakul, Kunat, et al.
Published: (2025)
by: Pipatanakul, Kunat, et al.
Published: (2025)
Typhoon T1: An Open Thai Reasoning Model
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
Prior Prompt Engineering for Reinforcement Fine-Tuning
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
by: Raina, Vyas, et al.
Published: (2024)
by: Raina, Vyas, et al.
Published: (2024)
Finetuning LLMs for Comparative Assessment Tasks
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
by: Chaichana, Yuatyong, et al.
Published: (2025)
by: Chaichana, Yuatyong, et al.
Published: (2025)
Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
by: Liusie, Adian, et al.
Published: (2024)
by: Liusie, Adian, et al.
Published: (2024)
WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
by: Molenda, Piotr, et al.
Published: (2024)
by: Molenda, Piotr, et al.
Published: (2024)
FinCoT: Grounding Chain-of-Thought in Expert Financial Reasoning
by: Nitarach, Natapong, et al.
Published: (2025)
by: Nitarach, Natapong, et al.
Published: (2025)
Investigating the Emergent Audio Classification Ability of ASR Foundation Models
by: Ma, Rao, et al.
Published: (2023)
by: Ma, Rao, et al.
Published: (2023)
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
by: Liusie, Adian, et al.
Published: (2024)
by: Liusie, Adian, et al.
Published: (2024)
Typhoon ASR Real-time: FastConformer-Transducer for Thai Automatic Speech Recognition
by: Sirichotedumrong, Warit, et al.
Published: (2026)
by: Sirichotedumrong, Warit, et al.
Published: (2026)
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
by: Sun, Guangzhi, et al.
Published: (2025)
by: Sun, Guangzhi, et al.
Published: (2025)
Mind the Gap! Static and Interactive Evaluations of Large Audio Models
by: Li, Minzhi, et al.
Published: (2025)
by: Li, Minzhi, et al.
Published: (2025)
Typhoon-S: Minimal Open Post-Training for Sovereign Large Language Models
by: Pipatanakul, Kunat, et al.
Published: (2026)
by: Pipatanakul, Kunat, et al.
Published: (2026)
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
by: Manakul, Potsawee, et al.
Published: (2026)
by: Manakul, Potsawee, et al.
Published: (2026)
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
by: Manakul, Potsawee, et al.
Published: (2025)
by: Manakul, Potsawee, et al.
Published: (2025)
Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models
by: Pipatanakul, Kunat, et al.
Published: (2024)
by: Pipatanakul, Kunat, et al.
Published: (2024)
Zero-shot Audio Topic Reranking using Large Language Models
by: Qian, Mengjie, et al.
Published: (2023)
by: Qian, Mengjie, et al.
Published: (2023)
Formula-One Prompting: A Composable Equation-First Prefix for Applied Mathematics
by: Nitarach, Natapong, et al.
Published: (2026)
by: Nitarach, Natapong, et al.
Published: (2026)
Cross-Lingual Interleaving for Speech Language Models
by: Moumen, Adel, et al.
Published: (2025)
by: Moumen, Adel, et al.
Published: (2025)
ThaiSafetyBench: Assessing Language Model Safety in Thai Cultural Contexts
by: Ukarapol, Trapoom, et al.
Published: (2026)
by: Ukarapol, Trapoom, et al.
Published: (2026)
Talk Less, Call Right: Enhancing Role-Play LLM Agents with Automatic Prompt Optimization and Role Prompting
by: Ruangtanusak, Saksorn, et al.
Published: (2025)
by: Ruangtanusak, Saksorn, et al.
Published: (2025)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
by: Sun, Guangzhi, et al.
Published: (2022)
by: Sun, Guangzhi, et al.
Published: (2022)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
by: Raina, Vyas, et al.
Published: (2024)
by: Raina, Vyas, et al.
Published: (2024)
Typhoon OCR: Open Vision-Language Model For Thai Document Extraction
by: Nonesung, Surapon, et al.
Published: (2026)
by: Nonesung, Surapon, et al.
Published: (2026)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Measuring the Redundancy of Decoder Layers in SpeechLLMs
by: Moumen, Adel, et al.
Published: (2026)
by: Moumen, Adel, et al.
Published: (2026)
On the Robustness of Answer Formats in Medical Reasoning Models
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
by: Zhao, Qiuming, et al.
Published: (2025)
by: Zhao, Qiuming, et al.
Published: (2025)
Mangosteen: An Open Thai Corpus for Language Model Pretraining
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
by: Lu, Xiaoding, et al.
Published: (2024)
by: Lu, Xiaoding, et al.
Published: (2024)
CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
by: Sun, Guangzhi, et al.
Published: (2025)
by: Sun, Guangzhi, et al.
Published: (2025)
Towards Better Understanding of Program-of-Thought Reasoning in Cross-Lingual and Multilingual Environments
by: Payoungkhamdee, Patomporn, et al.
Published: (2025)
by: Payoungkhamdee, Patomporn, et al.
Published: (2025)
Who can we trust? LLM-as-a-jury for Comparative Assessment
by: Qian, Mengjie, et al.
Published: (2026)
by: Qian, Mengjie, et al.
Published: (2026)
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
Protecting Bystander Privacy via Selective Hearing in Audio LLMs
by: Zhan, Xiao, et al.
Published: (2025)
by: Zhan, Xiao, et al.
Published: (2025)
Similar Items
-
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
by: Liusie, Adian, et al.
Published: (2023) -
SkillAggregation: Reference-free LLM-Dependent Aggregation
by: Sun, Guangzhi, et al.
Published: (2024) -
Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models
by: Manakul, Potsawee, et al.
Published: (2024) -
Adapting Language-Specific LLMs to a Reasoning Model in One Day via Model Merging -- An Open Recipe
by: Pipatanakul, Kunat, et al.
Published: (2025) -
Typhoon T1: An Open Thai Reasoning Model
by: Taveekitworachai, Pittawat, et al.
Published: (2025)