Winning Amazon KDD Cup'24
Fuente:
arXiv
Salvato in:
| Autori principali: | Deotte, Chris, Sorokin, Ivan, Erdem, Ahmet, Schifferer, Benedikt, Titericz Jr, Gilberto, Jegou, Simon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
di: Moshkov, Ivan, et al.
Pubblicazione: (2025)
di: Moshkov, Ivan, et al.
Pubblicazione: (2025)
DB3 Team's Solution For Meta KDD Cup' 25
di: Xia, Yikuan, et al.
Pubblicazione: (2025)
di: Xia, Yikuan, et al.
Pubblicazione: (2025)
KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
di: Jegou, Simon, et al.
Pubblicazione: (2026)
di: Jegou, Simon, et al.
Pubblicazione: (2026)
Neutral Residues: Revisiting Adapters for Model Extension
di: Talla, Franck Signe, et al.
Pubblicazione: (2024)
di: Talla, Franck Signe, et al.
Pubblicazione: (2024)
Revisiting the Solution of Meta KDD Cup 2024: CRAG
di: Ouyang, Jie, et al.
Pubblicazione: (2024)
di: Ouyang, Jie, et al.
Pubblicazione: (2024)
Mini-Giants: "Small" Language Models and Open Source Win-Win
di: Zhou, Zhengping, et al.
Pubblicazione: (2023)
di: Zhou, Zhengping, et al.
Pubblicazione: (2023)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
di: Han, Steve, et al.
Pubblicazione: (2025)
di: Han, Steve, et al.
Pubblicazione: (2025)
ToolACE: Winning the Points of LLM Function Calling
di: Liu, Weiwen, et al.
Pubblicazione: (2024)
di: Liu, Weiwen, et al.
Pubblicazione: (2024)
Random Masking Finds Winning Tickets for Parameter Efficient Fine-tuning
di: Xu, Jing, et al.
Pubblicazione: (2024)
di: Xu, Jing, et al.
Pubblicazione: (2024)
When Two LLMs Debate, Both Think They'll Win
di: Prasad, Pradyumna Shyama, et al.
Pubblicazione: (2025)
di: Prasad, Pradyumna Shyama, et al.
Pubblicazione: (2025)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
di: Chen, Sanxing, et al.
Pubblicazione: (2025)
di: Chen, Sanxing, et al.
Pubblicazione: (2025)
Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2024)
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2024)
Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents
di: Lewis, Ashley, et al.
Pubblicazione: (2025)
di: Lewis, Ashley, et al.
Pubblicazione: (2025)
Automated Factual Benchmarking for In-Car Conversational Systems using Large Language Models
di: Giebisch, Rafael, et al.
Pubblicazione: (2025)
di: Giebisch, Rafael, et al.
Pubblicazione: (2025)
MediaMind: Revolutionizing Media Monitoring using Agentification
di: Gunduz, Ahmet, et al.
Pubblicazione: (2025)
di: Gunduz, Ahmet, et al.
Pubblicazione: (2025)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
di: Husom, Erik Johannes, et al.
Pubblicazione: (2025)
di: Husom, Erik Johannes, et al.
Pubblicazione: (2025)
Data Science Kitchen at GermEval 2021: A Fine Selection of Hand-Picked Features, Delivered Fresh from the Oven
di: Hildebrandt, Niclas, et al.
Pubblicazione: (2021)
di: Hildebrandt, Niclas, et al.
Pubblicazione: (2021)
Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates
di: Zheng, Xiaosen, et al.
Pubblicazione: (2024)
di: Zheng, Xiaosen, et al.
Pubblicazione: (2024)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
di: Gritsch, Nikolas, et al.
Pubblicazione: (2024)
di: Gritsch, Nikolas, et al.
Pubblicazione: (2024)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
di: Devoto, Alessio, et al.
Pubblicazione: (2025)
di: Devoto, Alessio, et al.
Pubblicazione: (2025)
On Subjective Uncertainty Quantification and Calibration in Natural Language Generation
di: Wang, Ziyu, et al.
Pubblicazione: (2024)
di: Wang, Ziyu, et al.
Pubblicazione: (2024)
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions
di: Köksal, Abdullatif, et al.
Pubblicazione: (2024)
di: Köksal, Abdullatif, et al.
Pubblicazione: (2024)
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
di: Dang, John, et al.
Pubblicazione: (2024)
di: Dang, John, et al.
Pubblicazione: (2024)
Parameter Efficient Fine-tuning via Explained Variance Adaptation
di: Paischer, Fabian, et al.
Pubblicazione: (2024)
di: Paischer, Fabian, et al.
Pubblicazione: (2024)
Towards Efficient Neurally-Guided Program Induction for ARC-AGI
di: Ouellette, Simon
Pubblicazione: (2024)
di: Ouellette, Simon
Pubblicazione: (2024)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
di: Sainsbury, Chris, et al.
Pubblicazione: (2026)
di: Sainsbury, Chris, et al.
Pubblicazione: (2026)
Essential-Web v1.0: 24T tokens of organized web data
di: AI, Essential, et al.
Pubblicazione: (2025)
di: AI, Essential, et al.
Pubblicazione: (2025)
Reference-less Analysis of Context Specificity in Translation with Personalised Language Models
di: Vincent, Sebastian, et al.
Pubblicazione: (2023)
di: Vincent, Sebastian, et al.
Pubblicazione: (2023)
Observational Scaling Laws and the Predictability of Language Model Performance
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
Large Language Models Are Overparameterized Text Encoders
di: K, Thennal D, et al.
Pubblicazione: (2024)
di: K, Thennal D, et al.
Pubblicazione: (2024)
GLiREL -- Generalist Model for Zero-Shot Relation Extraction
di: Boylan, Jack, et al.
Pubblicazione: (2025)
di: Boylan, Jack, et al.
Pubblicazione: (2025)
Hidden State Poisoning Attacks against Mamba-based Language Models
di: Mercier, Alexandre Le, et al.
Pubblicazione: (2026)
di: Mercier, Alexandre Le, et al.
Pubblicazione: (2026)
What's New in My Data? Novelty Exploration via Contrastive Generation
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
On the Spatial Structure of Mixture-of-Experts in Transformers
di: Bershatsky, Daniel, et al.
Pubblicazione: (2025)
di: Bershatsky, Daniel, et al.
Pubblicazione: (2025)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
di: Mezentsev, Gleb, et al.
Pubblicazione: (2025)
di: Mezentsev, Gleb, et al.
Pubblicazione: (2025)
Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training
di: Kesgin, H. Toprak, et al.
Pubblicazione: (2024)
di: Kesgin, H. Toprak, et al.
Pubblicazione: (2024)
Discovering Forbidden Topics in Language Models
di: Rager, Can, et al.
Pubblicazione: (2025)
di: Rager, Can, et al.
Pubblicazione: (2025)
LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
di: Lugoloobi, William, et al.
Pubblicazione: (2026)
di: Lugoloobi, William, et al.
Pubblicazione: (2026)
Documenti analoghi
-
AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
di: Moshkov, Ivan, et al.
Pubblicazione: (2025) -
DB3 Team's Solution For Meta KDD Cup' 25
di: Xia, Yikuan, et al.
Pubblicazione: (2025) -
KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
di: Jegou, Simon, et al.
Pubblicazione: (2026) -
Neutral Residues: Revisiting Adapters for Model Extension
di: Talla, Franck Signe, et al.
Pubblicazione: (2024) -
Revisiting the Solution of Meta KDD Cup 2024: CRAG
di: Ouyang, Jie, et al.
Pubblicazione: (2024)