Efficient numeracy in language models through single-token number embeddings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kreitner, Linus, Hager, Paul, Mengedoht, Jonathan, Kaissis, Georgios, Rueckert, Daniel, Menten, Martin J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Survival In-Context: Amortized Bayesian Survival Analysis via Prior-Fitted Networks
von: Seletkov, Dmitrii, et al.
Veröffentlicht: (2026)
von: Seletkov, Dmitrii, et al.
Veröffentlicht: (2026)
A Tale of Two Classes: Adapting Supervised Contrastive Learning to Binary Imbalanced Datasets
von: Mildenberger, David, et al.
Veröffentlicht: (2025)
von: Mildenberger, David, et al.
Veröffentlicht: (2025)
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
von: Schwethelm, Kristian, et al.
Veröffentlicht: (2026)
von: Schwethelm, Kristian, et al.
Veröffentlicht: (2026)
Gradient-Weight Alignment as a Train-Time Proxy for Generalization in Classification Tasks
von: Hölzl, Florian A., et al.
Veröffentlicht: (2025)
von: Hölzl, Florian A., et al.
Veröffentlicht: (2025)
Incentivising the federation: gradient-based metrics for data selection and valuation in private decentralised training
von: Usynin, Dmitrii, et al.
Veröffentlicht: (2023)
von: Usynin, Dmitrii, et al.
Veröffentlicht: (2023)
On Arbitrary Predictions from Equally Valid Models
von: Lockfisch, Sarah, et al.
Veröffentlicht: (2025)
von: Lockfisch, Sarah, et al.
Veröffentlicht: (2025)
ChEX: Interactive Localization and Region Description in Chest X-rays
von: Müller, Philip, et al.
Veröffentlicht: (2024)
von: Müller, Philip, et al.
Veröffentlicht: (2024)
Laplace Sample Information: Data Informativeness Through a Bayesian Lens
von: Kaiser, Johannes, et al.
Veröffentlicht: (2025)
von: Kaiser, Johannes, et al.
Veröffentlicht: (2025)
Weakly Supervised Object Detection in Chest X-Rays with Differentiable ROI Proposal Networks and Soft ROI Pooling
von: Müller, Philip, et al.
Veröffentlicht: (2024)
von: Müller, Philip, et al.
Veröffentlicht: (2024)
Kernel Normalized Convolutional Networks
von: Nasirigerdeh, Reza, et al.
Veröffentlicht: (2022)
von: Nasirigerdeh, Reza, et al.
Veröffentlicht: (2022)
Weighting What Matters: Boosting Sample Efficiency in Medical Report Generation via Token Reweighting
von: Weers, Alexander, et al.
Veröffentlicht: (2026)
von: Weers, Alexander, et al.
Veröffentlicht: (2026)
Step-resolved data attribution for looped transformers
von: Kaissis, Georgios, et al.
Veröffentlicht: (2026)
von: Kaissis, Georgios, et al.
Veröffentlicht: (2026)
Unlocking the diagnostic potential of electrocardiograms through information transfer from cardiac magnetic resonance imaging
von: Turgut, Özgün, et al.
Veröffentlicht: (2023)
von: Turgut, Özgün, et al.
Veröffentlicht: (2023)
The cell as a token: high-dimensional geometry in language models and cell embeddings
von: Gilpin, William
Veröffentlicht: (2025)
von: Gilpin, William
Veröffentlicht: (2025)
Beyond the Calibration Point: Mechanism Comparison in Differential Privacy
von: Kaissis, Georgios, et al.
Veröffentlicht: (2024)
von: Kaissis, Georgios, et al.
Veröffentlicht: (2024)
Interpretable Retinal Disease Prediction Using Biology-Informed Heterogeneous Graph Representations
von: Lux, Laurin, et al.
Veröffentlicht: (2025)
von: Lux, Laurin, et al.
Veröffentlicht: (2025)
Improved Localized Machine Unlearning Through the Lens of Memorization
von: Torkzadehmahani, Reihaneh, et al.
Veröffentlicht: (2024)
von: Torkzadehmahani, Reihaneh, et al.
Veröffentlicht: (2024)
Machine Unlearning for Medical Imaging
von: Nasirigerdeh, Reza, et al.
Veröffentlicht: (2024)
von: Nasirigerdeh, Reza, et al.
Veröffentlicht: (2024)
Differentially Private Active Learning: Balancing Effective Data Selection and Privacy
von: Schwethelm, Kristian, et al.
Veröffentlicht: (2024)
von: Schwethelm, Kristian, et al.
Veröffentlicht: (2024)
Complex-valued Federated Learning with Differential Privacy and MRI Applications
von: Riess, Anneliese, et al.
Veröffentlicht: (2021)
von: Riess, Anneliese, et al.
Veröffentlicht: (2021)
Towards Generalisable Time Series Understanding Across Domains
von: Turgut, Özgün, et al.
Veröffentlicht: (2024)
von: Turgut, Özgün, et al.
Veröffentlicht: (2024)
Visualizing token importance for black-box language models
von: Rauba, Paulius, et al.
Veröffentlicht: (2025)
von: Rauba, Paulius, et al.
Veröffentlicht: (2025)
Do language models plan ahead for future tokens?
von: Wu, Wilson, et al.
Veröffentlicht: (2024)
von: Wu, Wilson, et al.
Veröffentlicht: (2024)
Unintended Memorization of Sensitive Information in Fine-Tuned Language Models
von: Szep, Marton, et al.
Veröffentlicht: (2026)
von: Szep, Marton, et al.
Veröffentlicht: (2026)
Private, fair and accurate: Training large-scale, privacy-preserving AI models in medical imaging
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2023)
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2023)
Long-range gene expression prediction with token alignment of large language model
von: Honig, Edouardo, et al.
Veröffentlicht: (2024)
von: Honig, Edouardo, et al.
Veröffentlicht: (2024)
Your Privacy Depends on Others: Collusion Vulnerabilities in Individual Differential Privacy
von: Kaiser, Johannes, et al.
Veröffentlicht: (2026)
von: Kaiser, Johannes, et al.
Veröffentlicht: (2026)
The Hitchhiker's Guide to Efficient, End-to-End, and Tight DP Auditing
von: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Veröffentlicht: (2025)
von: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Veröffentlicht: (2025)
Diff-Def: Diffusion-Generated Deformation Fields for Conditional Atlases
von: Starck, Sophie, et al.
Veröffentlicht: (2024)
von: Starck, Sophie, et al.
Veröffentlicht: (2024)
How Low Can You Go? Surfacing Prototypical In-Distribution Samples for Unsupervised Anomaly Detection
von: Meissen, Felix, et al.
Veröffentlicht: (2023)
von: Meissen, Felix, et al.
Veröffentlicht: (2023)
Beyond Benchmarks: Dynamic, Automatic And Systematic Red-Teaming Agents For Trustworthy Medical Language Models
von: Pan, Jiazhen, et al.
Veröffentlicht: (2025)
von: Pan, Jiazhen, et al.
Veröffentlicht: (2025)
Unified token representations for sequential decision models
von: Tian, Zhuojing, et al.
Veröffentlicht: (2025)
von: Tian, Zhuojing, et al.
Veröffentlicht: (2025)
LookupViT: Compressing visual information to a limited number of tokens
von: Koner, Rajat, et al.
Veröffentlicht: (2024)
von: Koner, Rajat, et al.
Veröffentlicht: (2024)
Pitfalls of topology-aware image segmentation
von: Berger, Alexander H., et al.
Veröffentlicht: (2024)
von: Berger, Alexander H., et al.
Veröffentlicht: (2024)
Interpretable deformable image registration: A geometric deep learning perspective
von: Sideri-Lampretsa, Vasiliki, et al.
Veröffentlicht: (2024)
von: Sideri-Lampretsa, Vasiliki, et al.
Veröffentlicht: (2024)
Not all tokens are needed(NAT): token efficient reinforcement learning
von: Sang, Hejian, et al.
Veröffentlicht: (2026)
von: Sang, Hejian, et al.
Veröffentlicht: (2026)
Learn from your own latents and not from tokens: A sample-complexity theory
von: Korchinski, Daniel J., et al.
Veröffentlicht: (2026)
von: Korchinski, Daniel J., et al.
Veröffentlicht: (2026)
Next-token pretraining implies in-context learning
von: Riechers, Paul M., et al.
Veröffentlicht: (2025)
von: Riechers, Paul M., et al.
Veröffentlicht: (2025)
Counterfactual Influence as a Distributional Quantity
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
Is your algorithm unlearning or untraining?
von: Triantafillou, Eleni, et al.
Veröffentlicht: (2026)
von: Triantafillou, Eleni, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Survival In-Context: Amortized Bayesian Survival Analysis via Prior-Fitted Networks
von: Seletkov, Dmitrii, et al.
Veröffentlicht: (2026) -
A Tale of Two Classes: Adapting Supervised Contrastive Learning to Binary Imbalanced Datasets
von: Mildenberger, David, et al.
Veröffentlicht: (2025) -
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
von: Schwethelm, Kristian, et al.
Veröffentlicht: (2026) -
Gradient-Weight Alignment as a Train-Time Proxy for Generalization in Classification Tasks
von: Hölzl, Florian A., et al.
Veröffentlicht: (2025) -
Incentivising the federation: gradient-based metrics for data selection and valuation in private decentralised training
von: Usynin, Dmitrii, et al.
Veröffentlicht: (2023)