HDT: Hierarchical Document Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | He, Haoyu, Flicke, Markus, Buchmann, Jan, Gurevych, Iryna, Geiger, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HDT: Hierarchical Discrete Transformer for Multivariate Time Series Forecasting
by: Feng, Shibo, et al.
Published: (2025)
by: Feng, Shibo, et al.
Published: (2025)
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
by: Rohweder, Jonas, et al.
Published: (2026)
by: Rohweder, Jonas, et al.
Published: (2026)
Citation Failure: Definition, Analysis and Efficient Mitigation
by: Buchmann, Jan, et al.
Published: (2025)
by: Buchmann, Jan, et al.
Published: (2025)
Attribute or Abstain: Large Language Models as Long Document Assistants
by: Buchmann, Jan, et al.
Published: (2024)
by: Buchmann, Jan, et al.
Published: (2024)
Document Structure in Long Document Transformers
by: Buchmann, Jan, et al.
Published: (2024)
by: Buchmann, Jan, et al.
Published: (2024)
Early Period of Training Impacts Adaptation for Out-of-Distribution Generalization: An Empirical Study
by: Liu, Chen Cecilia, et al.
Published: (2024)
by: Liu, Chen Cecilia, et al.
Published: (2024)
MDPO: Overcoming the Training-Inference Divide of Masked Diffusion Language Models
by: He, Haoyu, et al.
Published: (2025)
by: He, Haoyu, et al.
Published: (2025)
ABCD-LINK: Annotation Bootstrapping for Cross-Document Fine-Grained Links
by: Basch, Serwar, et al.
Published: (2025)
by: Basch, Serwar, et al.
Published: (2025)
Auditing Language Model Unlearning via Information Decomposition
by: Goel, Anmol, et al.
Published: (2026)
by: Goel, Anmol, et al.
Published: (2026)
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring
by: Paul, Indraneil, et al.
Published: (2026)
by: Paul, Indraneil, et al.
Published: (2026)
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
by: Tamoyan, Hovhannes, et al.
Published: (2025)
by: Tamoyan, Hovhannes, et al.
Published: (2025)
Differentially Private Steering for Large Language Model Alignment
by: Goel, Anmol, et al.
Published: (2025)
by: Goel, Anmol, et al.
Published: (2025)
How Quantization Shapes Bias in Large Language Models
by: Marcuzzi, Federico, et al.
Published: (2025)
by: Marcuzzi, Federico, et al.
Published: (2025)
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2024)
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2024)
SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward Modelling
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2025)
by: Rizvi, Md Imbesat Hassan, et al.
Published: (2025)
Uncertainty-Aware Decoding with Minimum Bayes Risk
by: Daheim, Nico, et al.
Published: (2025)
by: Daheim, Nico, et al.
Published: (2025)
Towards Automated Error Discovery: A Study in Conversational AI
by: Petrak, Dominic, et al.
Published: (2025)
by: Petrak, Dominic, et al.
Published: (2025)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
by: Geng, Jiahui, et al.
Published: (2025)
by: Geng, Jiahui, et al.
Published: (2025)
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
by: Daheim, Nico, et al.
Published: (2024)
by: Daheim, Nico, et al.
Published: (2024)
An Efficient Quantum Classifier Based on Hamiltonian Representations
by: Tiblias, Federico, et al.
Published: (2025)
by: Tiblias, Federico, et al.
Published: (2025)
How to Weight Multitask Finetuning? Fast Previews via Bayesian Model-Merging
by: Maldonado, Hugo Monzón, et al.
Published: (2024)
by: Maldonado, Hugo Monzón, et al.
Published: (2024)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2026)
by: Orel, Daniil, et al.
Published: (2026)
An Invitation to Deep Reinforcement Learning
by: Jaeger, Bernhard, et al.
Published: (2023)
by: Jaeger, Bernhard, et al.
Published: (2023)
MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
by: Macina, Jakub, et al.
Published: (2025)
by: Macina, Jakub, et al.
Published: (2025)
Model Merging by Uncertainty-Based Gradient Matching
by: Daheim, Nico, et al.
Published: (2023)
by: Daheim, Nico, et al.
Published: (2023)
UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation
by: von Rad, Jonathan, et al.
Published: (2026)
by: von Rad, Jonathan, et al.
Published: (2026)
GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers
by: Miyato, Takeru, et al.
Published: (2023)
by: Miyato, Takeru, et al.
Published: (2023)
Pretrained Mobility Transformer: A Foundation Model for Human Mobility
by: Wu, Xinhua, et al.
Published: (2024)
by: Wu, Xinhua, et al.
Published: (2024)
Federated Learning in Genetics: Extended Analysis of Accuracy, Performance and Privacy Trade-offs
by: Hannemann, Anika, et al.
Published: (2024)
by: Hannemann, Anika, et al.
Published: (2024)
How Do Transformers Learn Variable Binding in Symbolic Programs?
by: Wu, Yiwei, et al.
Published: (2025)
by: Wu, Yiwei, et al.
Published: (2025)
RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobility
by: He, Haoyu, et al.
Published: (2025)
by: He, Haoyu, et al.
Published: (2025)
Hierarchical Transformer for Electrocardiogram Diagnosis
by: Tang, Xiaoya, et al.
Published: (2024)
by: Tang, Xiaoya, et al.
Published: (2024)
Artificial Kuramoto Oscillatory Neurons
by: Miyato, Takeru, et al.
Published: (2024)
by: Miyato, Takeru, et al.
Published: (2024)
Operational early warning of thunderstorm-driven power outages from open data: a two-stage machine learning approach
by: Stanishevska, Iryna, et al.
Published: (2025)
by: Stanishevska, Iryna, et al.
Published: (2025)
On the Role of Priors in Bayesian Causal Learning
by: Geiger, Bernhard C., et al.
Published: (2025)
by: Geiger, Bernhard C., et al.
Published: (2025)
Information Plane Analysis of Binary Neural Networks
by: Nothnagel, Maximilian, et al.
Published: (2026)
by: Nothnagel, Maximilian, et al.
Published: (2026)
Robustness and Regularization in Hierarchical Re-Basin
by: Franke, Benedikt, et al.
Published: (2025)
by: Franke, Benedikt, et al.
Published: (2025)
Data-independent Module-aware Pruning for Hierarchical Vision Transformers
by: He, Yang, et al.
Published: (2024)
by: He, Yang, et al.
Published: (2024)
HMT: Hierarchical Memory Transformer for Efficient Long Context Language Processing
by: He, Zifan, et al.
Published: (2024)
by: He, Zifan, et al.
Published: (2024)
FIRE: Fact-checking with Iterative Retrieval and Verification
by: Xie, Zhuohan, et al.
Published: (2024)
by: Xie, Zhuohan, et al.
Published: (2024)
Similar Items
-
HDT: Hierarchical Discrete Transformer for Multivariate Time Series Forecasting
by: Feng, Shibo, et al.
Published: (2025) -
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
by: Rohweder, Jonas, et al.
Published: (2026) -
Citation Failure: Definition, Analysis and Efficient Mitigation
by: Buchmann, Jan, et al.
Published: (2025) -
Attribute or Abstain: Large Language Models as Long Document Assistants
by: Buchmann, Jan, et al.
Published: (2024) -
Document Structure in Long Document Transformers
by: Buchmann, Jan, et al.
Published: (2024)