A Benchmark Construction and Evaluation Framework for Specialist Domains: Case Study on Defense-related Documents
Fuente:
arXiv
Guardado en:
| Autores principales: | Doan, Bao Gia, Joshi, Aditya, Elinas, Pantelis, Bodhankar, Aarya, Leslie, Oscar, Marchant, Tom, Salim, Flora |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
What am I missing here?: Evaluating Large Language Models for Masked Sentence Prediction
por: Wyatt, Charlie, et al.
Publicado: (2025)
por: Wyatt, Charlie, et al.
Publicado: (2025)
Spectraformer: A Unified Random Feature Framework for Transformer
por: Nguyen, Duke, et al.
Publicado: (2024)
por: Nguyen, Duke, et al.
Publicado: (2024)
Harnessing Test-time Adaptation for NLU tasks Involving Dialects of English
por: Nguyen, Duke, et al.
Publicado: (2025)
por: Nguyen, Duke, et al.
Publicado: (2025)
Alternatives To Next Token Prediction In Text Generation -- A Survey
por: Wyatt, Charlie, et al.
Publicado: (2025)
por: Wyatt, Charlie, et al.
Publicado: (2025)
SDE-Attention: Latent Attention in SDE-RNNs for Irregularly Sampled Time Series with Missing Data
por: Fang, Yuting, et al.
Publicado: (2025)
por: Fang, Yuting, et al.
Publicado: (2025)
Permutation-based Inference for Variational Learning of Directed Acyclic Graphs
por: Bonilla, Edwin V., et al.
Publicado: (2024)
por: Bonilla, Edwin V., et al.
Publicado: (2024)
Hybrid Deep Learning Model for Multiple Cache Side Channel Attacks Detection: A Comparative Analysis
por: Joshi, Tejal, et al.
Publicado: (2025)
por: Joshi, Tejal, et al.
Publicado: (2025)
Mask-Conditioned Voxel Diffusion for Joint Geometry and Color Inpainting
por: Sumuk, Aarya
Publicado: (2026)
por: Sumuk, Aarya
Publicado: (2026)
Massive-STEPS: Massive Semantic Trajectories for Understanding POI Check-ins -- Dataset and Benchmarks
por: Wongso, Wilson, et al.
Publicado: (2025)
por: Wongso, Wilson, et al.
Publicado: (2025)
Evaluating the Influences of Explanation Style on Human-AI Reliance
por: Casolin, Emma, et al.
Publicado: (2024)
por: Casolin, Emma, et al.
Publicado: (2024)
AuditNet: A Conversational AI-based Security Assistant [DEMO]
por: Deldari, Shohreh, et al.
Publicado: (2024)
por: Deldari, Shohreh, et al.
Publicado: (2024)
Adversarial Graph Neural Network Benchmarks: Towards Practical and Fair Evaluation
por: Ngo, Tran Gia Bao, et al.
Publicado: (2026)
por: Ngo, Tran Gia Bao, et al.
Publicado: (2026)
A Two-stage Transformer Framework for Temporal Localization of Distracted Driver Behaviors
por: Doan, Gia-Bao, et al.
Publicado: (2026)
por: Doan, Gia-Bao, et al.
Publicado: (2026)
There is No "apple" in Timeseries: Rethinking TSFM through the Lens of Invariance
por: Prabowo, Arian, et al.
Publicado: (2025)
por: Prabowo, Arian, et al.
Publicado: (2025)
DRIFT-Net: A Spectral--Coupled Neural Operator for PDEs Learning
por: Li, Jiayi, et al.
Publicado: (2025)
por: Li, Jiayi, et al.
Publicado: (2025)
RCAEval: A Benchmark for Root Cause Analysis of Microservice Systems with Telemetry Data
por: Pham, Luan, et al.
Publicado: (2024)
por: Pham, Luan, et al.
Publicado: (2024)
Design and Implementation of a Java-Based Client-Server Application
por: Patil, Omkar, et al.
Publicado: (2024)
por: Patil, Omkar, et al.
Publicado: (2024)
Discrete Time Crystal in quantum Sherrington-Kirkpatrick model
por: Bothra, Aarya, et al.
Publicado: (2025)
por: Bothra, Aarya, et al.
Publicado: (2025)
GHGbench: A Unified Multi-Entity, Multi-Task Benchmark for Carbon Emission Prediction
por: Duan, Yifan, et al.
Publicado: (2026)
por: Duan, Yifan, et al.
Publicado: (2026)
SOCIA-EVO: Automated Simulator Construction via Dual-Anchored Bi-Level Optimization
por: Hua, Yuncheng, et al.
Publicado: (2026)
por: Hua, Yuncheng, et al.
Publicado: (2026)
ViLCo-Bench: VIdeo Language COntinual learning Benchmark
por: Tang, Tianqi, et al.
Publicado: (2024)
por: Tang, Tianqi, et al.
Publicado: (2024)
On the Induced Neighbourhood of Vertex-Transitive Graphs
por: Joshi, Aditya
Publicado: (2025)
por: Joshi, Aditya
Publicado: (2025)
Code Interviews: Design and Evaluation of a More Authentic Assessment for Introductory Programming Assignments
por: Kannam, Suhas, et al.
Publicado: (2024)
por: Kannam, Suhas, et al.
Publicado: (2024)
Identifying high resolution benchmark data needs and Novel data-driven methodologies for Climate Downscaling
por: Curran, Declan, et al.
Publicado: (2024)
por: Curran, Declan, et al.
Publicado: (2024)
STC-ViT: Spatio Temporal Continuous Vision Transformer for Medium-range Global Weather Forecasting
por: Saleem, Hira, et al.
Publicado: (2024)
por: Saleem, Hira, et al.
Publicado: (2024)
Illusions of reflection: open-ended task reveals systematic failures in Large Language Models' reflective reasoning
por: Weatherhead, Sion, et al.
Publicado: (2025)
por: Weatherhead, Sion, et al.
Publicado: (2025)
A-UTE: Advection Informed, Uncertainty Aware Temperature Emulator
por: Saleem, Hira, et al.
Publicado: (2024)
por: Saleem, Hira, et al.
Publicado: (2024)
PINN-Cast: Exploring the Role of Continuous-Depth NODE in Transformers and Physics Informed Loss as Soft Physical Constraints in Short-term Weather Forecasting
por: Saleem, Hira, et al.
Publicado: (2026)
por: Saleem, Hira, et al.
Publicado: (2026)
Mechanistic Indicators of Steering Effectiveness in Large Language Models
por: Jafari, Mehdi, et al.
Publicado: (2026)
por: Jafari, Mehdi, et al.
Publicado: (2026)
Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
por: Bui, Khanh Gia
Publicado: (2025)
por: Bui, Khanh Gia
Publicado: (2025)
Multi-Stage Verification-Centric Framework for Mitigating Hallucination in Multi-Modal RAG
por: Chen, Baiyu, et al.
Publicado: (2025)
por: Chen, Baiyu, et al.
Publicado: (2025)
Colombia – US relations in an era of great power competition
por: Aaron Marchant, et al.
Publicado: (2024)
por: Aaron Marchant, et al.
Publicado: (2024)
Can Current AI Models Count What We Mean, Not What They See? A Benchmark and Systematic Evaluation
por: Nguyen, Gia Khanh, et al.
Publicado: (2025)
por: Nguyen, Gia Khanh, et al.
Publicado: (2025)
DeepStage: Learning Autonomous Defense Policies Against Multi-Stage APT Campaigns
por: Phan, Trung V., et al.
Publicado: (2026)
por: Phan, Trung V., et al.
Publicado: (2026)
HiT-JEPA: A Hierarchical Self-supervised Trajectory Embedding Framework for Similarity Computation
por: Li, Lihuan, et al.
Publicado: (2025)
por: Li, Lihuan, et al.
Publicado: (2025)
TrajLLM: A Modular LLM-Enhanced Agent-Based Framework for Realistic Human Trajectory Simulation
por: Ju, Chenlu, et al.
Publicado: (2025)
por: Ju, Chenlu, et al.
Publicado: (2025)
ODEStream: A Buffer-Free Online Learning Framework with ODE-based Adaptor for Streaming Time Series Forecasting
por: Abushaqra, Futoon M., et al.
Publicado: (2024)
por: Abushaqra, Futoon M., et al.
Publicado: (2024)
Understanding Structural Dynamics in Flexible Rare‐Earth Metal‐Organic Frameworks
por: Edward Loukopoulos, et al.
Publicado: (2024)
por: Edward Loukopoulos, et al.
Publicado: (2024)
AutoB2G: A Large Language Model-Driven Agentic Framework For Automated Building-Grid Co-Simulation
por: Zhang, Borui, et al.
Publicado: (2026)
por: Zhang, Borui, et al.
Publicado: (2026)
BESSTIE: A Benchmark for Sentiment and Sarcasm Classification for Varieties of English
por: Srirag, Dipankar, et al.
Publicado: (2024)
por: Srirag, Dipankar, et al.
Publicado: (2024)
Ejemplares similares
-
What am I missing here?: Evaluating Large Language Models for Masked Sentence Prediction
por: Wyatt, Charlie, et al.
Publicado: (2025) -
Spectraformer: A Unified Random Feature Framework for Transformer
por: Nguyen, Duke, et al.
Publicado: (2024) -
Harnessing Test-time Adaptation for NLU tasks Involving Dialects of English
por: Nguyen, Duke, et al.
Publicado: (2025) -
Alternatives To Next Token Prediction In Text Generation -- A Survey
por: Wyatt, Charlie, et al.
Publicado: (2025) -
SDE-Attention: Latent Attention in SDE-RNNs for Irregularly Sampled Time Series with Missing Data
por: Fang, Yuting, et al.
Publicado: (2025)