AuthTrace: Diagnosing Evidence Construction in Thematically Dense Single-Author Corpora
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Xiaoqing, Li, Feifei, Ming, Haoliang, Que, Wenhui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Retrieval as Reasoning: Self-Evolving Agent-Native Retrieval via LLM-Wiki
por: Ming, Haoliang, et al.
Publicado: (2026)
por: Ming, Haoliang, et al.
Publicado: (2026)
MemCog: From Memory-as-Tool to Memory-as-Cognition in Conversational Agents
por: Li, Zihan, et al.
Publicado: (2026)
por: Li, Zihan, et al.
Publicado: (2026)
Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses
por: Peng, Kerui, et al.
Publicado: (2026)
por: Peng, Kerui, et al.
Publicado: (2026)
One Agent to Serve All: a Lite-Adaptive Stylized AI Assistant for Millions of Multi-Style Official Accounts
por: Fan, Xingyu, et al.
Publicado: (2025)
por: Fan, Xingyu, et al.
Publicado: (2025)
BN-AuthProf: Benchmarking Machine Learning for Bangla Author Profiling on Social Media Texts
por: Tasnim, Raisa, et al.
Publicado: (2024)
por: Tasnim, Raisa, et al.
Publicado: (2024)
Bidirectional Topic Matching: Quantifying Thematic Overlap Between Corpora Through Topic Modelling
por: Adam, Raven, et al.
Publicado: (2024)
por: Adam, Raven, et al.
Publicado: (2024)
When Should Dense Retrievers Be Updated in Evolving Corpora? Detecting Out-of-Distribution Corpora Using GradNormIR
por: Ko, Dayoon, et al.
Publicado: (2025)
por: Ko, Dayoon, et al.
Publicado: (2025)
MASRAD: Arabic Terminology Management Corpora with Semi-Automatic Construction
por: Nasser, Mahdi, et al.
Publicado: (2025)
por: Nasser, Mahdi, et al.
Publicado: (2025)
Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification
por: Qiu, Jingxi, et al.
Publicado: (2026)
por: Qiu, Jingxi, et al.
Publicado: (2026)
A Hierarchical and Attentional Analysis of Argument Structure Constructions in BERT Using Naturalistic Corpora
por: Kaipeng, Liu, et al.
Publicado: (2026)
por: Kaipeng, Liu, et al.
Publicado: (2026)
Building Corpora for Single-Channel Speech Separation Across Multiple Domains
por: Maciejewski, Matthew, et al.
Publicado: (2018)
por: Maciejewski, Matthew, et al.
Publicado: (2018)
MuCPT: Music-related Natural Language Model Continued Pretraining
por: Tian, Kai, et al.
Publicado: (2025)
por: Tian, Kai, et al.
Publicado: (2025)
Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora
por: Hennara, Khalil, et al.
Publicado: (2025)
por: Hennara, Khalil, et al.
Publicado: (2025)
A Method for Learning Large-Scale Computational Construction Grammars from Semantically Annotated Corpora
por: Van Eecke, Paul, et al.
Publicado: (2026)
por: Van Eecke, Paul, et al.
Publicado: (2026)
Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora
por: Abbas, Chaymaa, et al.
Publicado: (2026)
por: Abbas, Chaymaa, et al.
Publicado: (2026)
Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs
por: Zhang, Jiaqiao, et al.
Publicado: (2026)
por: Zhang, Jiaqiao, et al.
Publicado: (2026)
TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research Corpora
por: Kargupta, Priyanka, et al.
Publicado: (2025)
por: Kargupta, Priyanka, et al.
Publicado: (2025)
Validating and Exploring Large Geographic Corpora
por: Dunn, Jonathan
Publicado: (2024)
por: Dunn, Jonathan
Publicado: (2024)
Identifying Emerging Concepts in Large Corpora
por: Ma, Sibo, et al.
Publicado: (2025)
por: Ma, Sibo, et al.
Publicado: (2025)
Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs
por: Han, Peitao, et al.
Publicado: (2026)
por: Han, Peitao, et al.
Publicado: (2026)
Dense SAE Latents Are Features, Not Bugs
por: Sun, Xiaoqing, et al.
Publicado: (2025)
por: Sun, Xiaoqing, et al.
Publicado: (2025)
Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge
por: Wu, Xuanxin, et al.
Publicado: (2025)
por: Wu, Xuanxin, et al.
Publicado: (2025)
AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora
por: Bai, Jiaxin, et al.
Publicado: (2025)
por: Bai, Jiaxin, et al.
Publicado: (2025)
Reflections on Inductive Thematic Saturation as a potential metric for measuring the validity of an inductive Thematic Analysis with LLMs
por: De Paoli, Stefano, et al.
Publicado: (2024)
por: De Paoli, Stefano, et al.
Publicado: (2024)
MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards
por: Shen, Zhiyu, et al.
Publicado: (2026)
por: Shen, Zhiyu, et al.
Publicado: (2026)
New Textual Corpora for Serbian Language Modeling
por: Škorić, Mihailo, et al.
Publicado: (2024)
por: Škorić, Mihailo, et al.
Publicado: (2024)
Bias in News Summarization: Measures, Pitfalls and Corpora
por: Steen, Julius, et al.
Publicado: (2023)
por: Steen, Julius, et al.
Publicado: (2023)
Comparable Corpora: Opportunities for New Research Directions
por: Church, Kenneth
Publicado: (2025)
por: Church, Kenneth
Publicado: (2025)
Attributing Culture-Conditioned Generations to Pretraining Corpora
por: Li, Huihan, et al.
Publicado: (2024)
por: Li, Huihan, et al.
Publicado: (2024)
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
por: Pungeršek, Taja Kuzman, et al.
Publicado: (2026)
por: Pungeršek, Taja Kuzman, et al.
Publicado: (2026)
Mathematical Entities: Corpora and Benchmarks
por: Collard, Jacob, et al.
Publicado: (2024)
por: Collard, Jacob, et al.
Publicado: (2024)
InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning
por: Wei, Chengwei, et al.
Publicado: (2026)
por: Wei, Chengwei, et al.
Publicado: (2026)
Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction
por: Li, Zehan, et al.
Publicado: (2026)
por: Li, Zehan, et al.
Publicado: (2026)
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video
por: Que, Shumin, et al.
Publicado: (2025)
por: Que, Shumin, et al.
Publicado: (2025)
Deep Learning Based Dense Retrieval: A Comparative Study
por: Zhong, Ming, et al.
Publicado: (2024)
por: Zhong, Ming, et al.
Publicado: (2024)
AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
por: Milička, Jiří, et al.
Publicado: (2025)
por: Milička, Jiří, et al.
Publicado: (2025)
Wiki Dumps to Training Corpora: South Slavic Case
por: Škorić, Mihailo, et al.
Publicado: (2026)
por: Škorić, Mihailo, et al.
Publicado: (2026)
Disambiguating Numeral Sequences to Decipher Ancient Accounting Corpora
por: Born, Logan, et al.
Publicado: (2025)
por: Born, Logan, et al.
Publicado: (2025)
Active Learning for Multilingual Fingerspelling Corpora
por: Wang, Shuai, et al.
Publicado: (2023)
por: Wang, Shuai, et al.
Publicado: (2023)
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
por: Fei, Zhaoye, et al.
Publicado: (2024)
por: Fei, Zhaoye, et al.
Publicado: (2024)
Ejemplares similares
-
Retrieval as Reasoning: Self-Evolving Agent-Native Retrieval via LLM-Wiki
por: Ming, Haoliang, et al.
Publicado: (2026) -
MemCog: From Memory-as-Tool to Memory-as-Cognition in Conversational Agents
por: Li, Zihan, et al.
Publicado: (2026) -
Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses
por: Peng, Kerui, et al.
Publicado: (2026) -
One Agent to Serve All: a Lite-Adaptive Stylized AI Assistant for Millions of Multi-Style Official Accounts
por: Fan, Xingyu, et al.
Publicado: (2025) -
BN-AuthProf: Benchmarking Machine Learning for Bangla Author Profiling on Social Media Texts
por: Tasnim, Raisa, et al.
Publicado: (2024)