Blocks Architecture (BloArk): Efficient, Cost-Effective, and Incremental Dataset Architecture for Wikipedia Revision History
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Lingxi, Yao, Zonghai, Kwon, Sunjae, Yu, Hong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
por: Smădu, Răzvan-Alexandru, et al.
Publicado: (2025)
por: Smădu, Răzvan-Alexandru, et al.
Publicado: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
por: Ashuach, Tomer, et al.
Publicado: (2025)
por: Ashuach, Tomer, et al.
Publicado: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
por: Collado-Montañez, Jaime, et al.
Publicado: (2025)
por: Collado-Montañez, Jaime, et al.
Publicado: (2025)
Graphemic Normalization of the Perso-Arabic Script
por: Doctor, Raiomond, et al.
Publicado: (2022)
por: Doctor, Raiomond, et al.
Publicado: (2022)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
por: Gutkin, Alexander, et al.
Publicado: (2023)
por: Gutkin, Alexander, et al.
Publicado: (2023)
SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language Models
por: Liu, Han, et al.
Publicado: (2026)
por: Liu, Han, et al.
Publicado: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
por: Souza, Débora, et al.
Publicado: (2026)
por: Souza, Débora, et al.
Publicado: (2026)
Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation
por: Bertina, Abbas, et al.
Publicado: (2025)
por: Bertina, Abbas, et al.
Publicado: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models
por: Wen, Zhihao, et al.
Publicado: (2025)
por: Wen, Zhihao, et al.
Publicado: (2025)
SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM
por: Wang, Quandong, et al.
Publicado: (2024)
por: Wang, Quandong, et al.
Publicado: (2024)
DimStance: Multilingual Datasets for Dimensional Stance Analysis
por: Becker, Jonas, et al.
Publicado: (2026)
por: Becker, Jonas, et al.
Publicado: (2026)
Efficient Knowledge Feeding to Language Models: A Novel Integrated Encoder-Decoder Architecture
por: Kumar, S Santosh, et al.
Publicado: (2025)
por: Kumar, S Santosh, et al.
Publicado: (2025)
DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
por: Lee, Lung-Hao, et al.
Publicado: (2026)
por: Lee, Lung-Hao, et al.
Publicado: (2026)
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
por: Hu, Yuxuan, et al.
Publicado: (2025)
por: Hu, Yuxuan, et al.
Publicado: (2025)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
Locations of Characters in Narratives: Andersen and Persuasion Datasets
por: Ozyurt, Batuhan, et al.
Publicado: (2025)
por: Ozyurt, Batuhan, et al.
Publicado: (2025)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
por: Toker, Michael, et al.
Publicado: (2024)
por: Toker, Michael, et al.
Publicado: (2024)
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin
por: Bothwell, Stephen, et al.
Publicado: (2024)
por: Bothwell, Stephen, et al.
Publicado: (2024)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
por: Bouchekif, Abdessalam, et al.
Publicado: (2026)
por: Bouchekif, Abdessalam, et al.
Publicado: (2026)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
ML-Promise: A Multilingual Dataset for Corporate Promise Verification
por: Seki, Yohei, et al.
Publicado: (2024)
por: Seki, Yohei, et al.
Publicado: (2024)
Breaking the HISCO Barrier: Automatic Occupational Standardization with OccCANINE
por: Dahl, Christian Møller, et al.
Publicado: (2024)
por: Dahl, Christian Møller, et al.
Publicado: (2024)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
por: Tu, Songjun, et al.
Publicado: (2026)
por: Tu, Songjun, et al.
Publicado: (2026)
Low-Resource Court Judgment Summarization for Common Law Systems
por: Liu, Shuaiqi, et al.
Publicado: (2024)
por: Liu, Shuaiqi, et al.
Publicado: (2024)
EMO-KNOW: A Large Scale Dataset on Emotion and Emotion-cause
por: Nguyen, Mia Huong, et al.
Publicado: (2024)
por: Nguyen, Mia Huong, et al.
Publicado: (2024)
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
por: Sileo, Damien
Publicado: (2024)
por: Sileo, Damien
Publicado: (2024)
Comparative Analysis of AI Agent Architectures for Entity Relationship Classification
por: Berijanian, Maryam, et al.
Publicado: (2025)
por: Berijanian, Maryam, et al.
Publicado: (2025)
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis
por: Bose, Joy
Publicado: (2026)
por: Bose, Joy
Publicado: (2026)
Tracking Semantic Change in Slovene: A Novel Dataset and Optimal Transport-Based Distance
por: Pranjić, Marko, et al.
Publicado: (2024)
por: Pranjić, Marko, et al.
Publicado: (2024)
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
por: Lequeu, Pierre-Antoine, et al.
Publicado: (2026)
por: Lequeu, Pierre-Antoine, et al.
Publicado: (2026)
AMELI: Enhancing Multimodal Entity Linking with Fine-Grained Attributes
por: Yao, Barry Menglong, et al.
Publicado: (2023)
por: Yao, Barry Menglong, et al.
Publicado: (2023)
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
por: Garg, Aashna, et al.
Publicado: (2026)
por: Garg, Aashna, et al.
Publicado: (2026)
I run as fast as a rabbit, can you? A Multilingual Simile Dialogue Dataset
por: Ma, Longxuan, et al.
Publicado: (2023)
por: Ma, Longxuan, et al.
Publicado: (2023)
ELECTRA and GPT-4o: Cost-Effective Partners for Sentiment Analysis
por: Beno, James P.
Publicado: (2024)
por: Beno, James P.
Publicado: (2024)
RUQuant: Towards Refining Uniform Quantization for Large Language Models
por: Liu, Han, et al.
Publicado: (2026)
por: Liu, Han, et al.
Publicado: (2026)
Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development
por: Sahai, Sattvik, et al.
Publicado: (2025)
por: Sahai, Sattvik, et al.
Publicado: (2025)
MIMIC-SR-ICD11: A Dataset for Narrative-Based Diagnosis
por: Wu, Yuexin, et al.
Publicado: (2025)
por: Wu, Yuexin, et al.
Publicado: (2025)
Pipeline and Dataset Generation for Automated Fact-checking in Almost Any Language
por: Drchal, Jan, et al.
Publicado: (2023)
por: Drchal, Jan, et al.
Publicado: (2023)
Ejemplares similares
-
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
por: Smădu, Răzvan-Alexandru, et al.
Publicado: (2025) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
por: Ashuach, Tomer, et al.
Publicado: (2025) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025) -
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
por: Collado-Montañez, Jaime, et al.
Publicado: (2025) -
Graphemic Normalization of the Perso-Arabic Script
por: Doctor, Raiomond, et al.
Publicado: (2022)