HDT: Hierarchical Document Transformer
Fuente:
arXiv
Guardado en:
| Autores principales: | He, Haoyu, Flicke, Markus, Buchmann, Jan, Gurevych, Iryna, Geiger, Andreas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HDT: Hierarchical Discrete Transformer for Multivariate Time Series Forecasting
por: Feng, Shibo, et al.
Publicado: (2025)
por: Feng, Shibo, et al.
Publicado: (2025)
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
por: Rohweder, Jonas, et al.
Publicado: (2026)
por: Rohweder, Jonas, et al.
Publicado: (2026)
Citation Failure: Definition, Analysis and Efficient Mitigation
por: Buchmann, Jan, et al.
Publicado: (2025)
por: Buchmann, Jan, et al.
Publicado: (2025)
Attribute or Abstain: Large Language Models as Long Document Assistants
por: Buchmann, Jan, et al.
Publicado: (2024)
por: Buchmann, Jan, et al.
Publicado: (2024)
Document Structure in Long Document Transformers
por: Buchmann, Jan, et al.
Publicado: (2024)
por: Buchmann, Jan, et al.
Publicado: (2024)
Early Period of Training Impacts Adaptation for Out-of-Distribution Generalization: An Empirical Study
por: Liu, Chen Cecilia, et al.
Publicado: (2024)
por: Liu, Chen Cecilia, et al.
Publicado: (2024)
MDPO: Overcoming the Training-Inference Divide of Masked Diffusion Language Models
por: He, Haoyu, et al.
Publicado: (2025)
por: He, Haoyu, et al.
Publicado: (2025)
ABCD-LINK: Annotation Bootstrapping for Cross-Document Fine-Grained Links
por: Basch, Serwar, et al.
Publicado: (2025)
por: Basch, Serwar, et al.
Publicado: (2025)
Auditing Language Model Unlearning via Information Decomposition
por: Goel, Anmol, et al.
Publicado: (2026)
por: Goel, Anmol, et al.
Publicado: (2026)
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring
por: Paul, Indraneil, et al.
Publicado: (2026)
por: Paul, Indraneil, et al.
Publicado: (2026)
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
por: Tamoyan, Hovhannes, et al.
Publicado: (2025)
por: Tamoyan, Hovhannes, et al.
Publicado: (2025)
Differentially Private Steering for Large Language Model Alignment
por: Goel, Anmol, et al.
Publicado: (2025)
por: Goel, Anmol, et al.
Publicado: (2025)
How Quantization Shapes Bias in Large Language Models
por: Marcuzzi, Federico, et al.
Publicado: (2025)
por: Marcuzzi, Federico, et al.
Publicado: (2025)
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
por: Rizvi, Md Imbesat Hassan, et al.
Publicado: (2024)
por: Rizvi, Md Imbesat Hassan, et al.
Publicado: (2024)
SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward Modelling
por: Rizvi, Md Imbesat Hassan, et al.
Publicado: (2025)
por: Rizvi, Md Imbesat Hassan, et al.
Publicado: (2025)
Uncertainty-Aware Decoding with Minimum Bayes Risk
por: Daheim, Nico, et al.
Publicado: (2025)
por: Daheim, Nico, et al.
Publicado: (2025)
Towards Automated Error Discovery: A Study in Conversational AI
por: Petrak, Dominic, et al.
Publicado: (2025)
por: Petrak, Dominic, et al.
Publicado: (2025)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
por: Geng, Jiahui, et al.
Publicado: (2025)
por: Geng, Jiahui, et al.
Publicado: (2025)
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
por: Daheim, Nico, et al.
Publicado: (2024)
por: Daheim, Nico, et al.
Publicado: (2024)
An Efficient Quantum Classifier Based on Hamiltonian Representations
por: Tiblias, Federico, et al.
Publicado: (2025)
por: Tiblias, Federico, et al.
Publicado: (2025)
How to Weight Multitask Finetuning? Fast Previews via Bayesian Model-Merging
por: Maldonado, Hugo Monzón, et al.
Publicado: (2024)
por: Maldonado, Hugo Monzón, et al.
Publicado: (2024)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
por: Orel, Daniil, et al.
Publicado: (2026)
por: Orel, Daniil, et al.
Publicado: (2026)
An Invitation to Deep Reinforcement Learning
por: Jaeger, Bernhard, et al.
Publicado: (2023)
por: Jaeger, Bernhard, et al.
Publicado: (2023)
MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
por: Macina, Jakub, et al.
Publicado: (2025)
por: Macina, Jakub, et al.
Publicado: (2025)
Model Merging by Uncertainty-Based Gradient Matching
por: Daheim, Nico, et al.
Publicado: (2023)
por: Daheim, Nico, et al.
Publicado: (2023)
UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation
por: von Rad, Jonathan, et al.
Publicado: (2026)
por: von Rad, Jonathan, et al.
Publicado: (2026)
GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers
por: Miyato, Takeru, et al.
Publicado: (2023)
por: Miyato, Takeru, et al.
Publicado: (2023)
Pretrained Mobility Transformer: A Foundation Model for Human Mobility
por: Wu, Xinhua, et al.
Publicado: (2024)
por: Wu, Xinhua, et al.
Publicado: (2024)
Federated Learning in Genetics: Extended Analysis of Accuracy, Performance and Privacy Trade-offs
por: Hannemann, Anika, et al.
Publicado: (2024)
por: Hannemann, Anika, et al.
Publicado: (2024)
How Do Transformers Learn Variable Binding in Symbolic Programs?
por: Wu, Yiwei, et al.
Publicado: (2025)
por: Wu, Yiwei, et al.
Publicado: (2025)
RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobility
por: He, Haoyu, et al.
Publicado: (2025)
por: He, Haoyu, et al.
Publicado: (2025)
Hierarchical Transformer for Electrocardiogram Diagnosis
por: Tang, Xiaoya, et al.
Publicado: (2024)
por: Tang, Xiaoya, et al.
Publicado: (2024)
Artificial Kuramoto Oscillatory Neurons
por: Miyato, Takeru, et al.
Publicado: (2024)
por: Miyato, Takeru, et al.
Publicado: (2024)
Operational early warning of thunderstorm-driven power outages from open data: a two-stage machine learning approach
por: Stanishevska, Iryna, et al.
Publicado: (2025)
por: Stanishevska, Iryna, et al.
Publicado: (2025)
On the Role of Priors in Bayesian Causal Learning
por: Geiger, Bernhard C., et al.
Publicado: (2025)
por: Geiger, Bernhard C., et al.
Publicado: (2025)
Information Plane Analysis of Binary Neural Networks
por: Nothnagel, Maximilian, et al.
Publicado: (2026)
por: Nothnagel, Maximilian, et al.
Publicado: (2026)
Robustness and Regularization in Hierarchical Re-Basin
por: Franke, Benedikt, et al.
Publicado: (2025)
por: Franke, Benedikt, et al.
Publicado: (2025)
Data-independent Module-aware Pruning for Hierarchical Vision Transformers
por: He, Yang, et al.
Publicado: (2024)
por: He, Yang, et al.
Publicado: (2024)
HMT: Hierarchical Memory Transformer for Efficient Long Context Language Processing
por: He, Zifan, et al.
Publicado: (2024)
por: He, Zifan, et al.
Publicado: (2024)
FIRE: Fact-checking with Iterative Retrieval and Verification
por: Xie, Zhuohan, et al.
Publicado: (2024)
por: Xie, Zhuohan, et al.
Publicado: (2024)
Ejemplares similares
-
HDT: Hierarchical Discrete Transformer for Multivariate Time Series Forecasting
por: Feng, Shibo, et al.
Publicado: (2025) -
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
por: Rohweder, Jonas, et al.
Publicado: (2026) -
Citation Failure: Definition, Analysis and Efficient Mitigation
por: Buchmann, Jan, et al.
Publicado: (2025) -
Attribute or Abstain: Large Language Models as Long Document Assistants
por: Buchmann, Jan, et al.
Publicado: (2024) -
Document Structure in Long Document Transformers
por: Buchmann, Jan, et al.
Publicado: (2024)