Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Santos, Matheus L. O., Campelo, Cláudio E. C. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
von: Teixeira, Tiago, et al.
Veröffentlicht: (2026)
von: Teixeira, Tiago, et al.
Veröffentlicht: (2026)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
von: Du, Bangde, et al.
Veröffentlicht: (2025)
von: Du, Bangde, et al.
Veröffentlicht: (2025)
ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration
von: Deng, Minghang, et al.
Veröffentlicht: (2025)
von: Deng, Minghang, et al.
Veröffentlicht: (2025)
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
von: Williamson, Dane, et al.
Veröffentlicht: (2025)
von: Williamson, Dane, et al.
Veröffentlicht: (2025)
LLMs and the Human Condition
von: Wallis, Peter
Veröffentlicht: (2024)
von: Wallis, Peter
Veröffentlicht: (2024)
ChemPro: A Progressive Chemistry Benchmark for Large Language Models
von: Baranwal, Aaditya, et al.
Veröffentlicht: (2026)
von: Baranwal, Aaditya, et al.
Veröffentlicht: (2026)
On Explaining with Attention Matrices
von: Naim, Omar, et al.
Veröffentlicht: (2024)
von: Naim, Omar, et al.
Veröffentlicht: (2024)
Text2Model: Modeling Copilots for Text-to-Model Translation
von: Kadioglu, Serdar, et al.
Veröffentlicht: (2026)
von: Kadioglu, Serdar, et al.
Veröffentlicht: (2026)
Quo Vadis ChatGPT? From Large Language Models to Large Knowledge Models
von: Venkatasubramanian, Venkat, et al.
Veröffentlicht: (2024)
von: Venkatasubramanian, Venkat, et al.
Veröffentlicht: (2024)
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax?
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2023)
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2023)
Automated Circuit Interpretation via Probe Prompting
von: Birardi, Giuseppe
Veröffentlicht: (2025)
von: Birardi, Giuseppe
Veröffentlicht: (2025)
Aspect-Based Sentiment Analysis for Future Tourism Experiences: A BERT-MoE Framework for Persian User Reviews
von: Taskooh, Hamidreza Kazemi, et al.
Veröffentlicht: (2026)
von: Taskooh, Hamidreza Kazemi, et al.
Veröffentlicht: (2026)
Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
von: Pareja, Aldo, et al.
Veröffentlicht: (2024)
von: Pareja, Aldo, et al.
Veröffentlicht: (2024)
A Survey of Text and Speech Resources for Hausa and Fongbe: Availability, Quality, and Gaps for NLP Development
von: Adjovi, Mahounan Pericles, et al.
Veröffentlicht: (2026)
von: Adjovi, Mahounan Pericles, et al.
Veröffentlicht: (2026)
Pareto-Optimized Open-Source LLMs for Healthcare via Context Retrieval
von: Bayarri-Planas, Jordi, et al.
Veröffentlicht: (2024)
von: Bayarri-Planas, Jordi, et al.
Veröffentlicht: (2024)
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2025)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2025)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
von: Goldstein, Daniel, et al.
Veröffentlicht: (2026)
von: Goldstein, Daniel, et al.
Veröffentlicht: (2026)
Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
von: Plevris, Vagelis, et al.
Veröffentlicht: (2023)
von: Plevris, Vagelis, et al.
Veröffentlicht: (2023)
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
von: Edin, Joakim, et al.
Veröffentlicht: (2025)
von: Edin, Joakim, et al.
Veröffentlicht: (2025)
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
von: Buchner, Valentin Leonhard, et al.
Veröffentlicht: (2023)
von: Buchner, Valentin Leonhard, et al.
Veröffentlicht: (2023)
NRR-Core: Non-Resolution Reasoning as a Computational Framework for Contextual Identity and Ambiguity Preservation
von: Saito, Kei
Veröffentlicht: (2025)
von: Saito, Kei
Veröffentlicht: (2025)
ALISON: Fast and Effective Stylometric Authorship Obfuscation
von: Xing, Eric, et al.
Veröffentlicht: (2024)
von: Xing, Eric, et al.
Veröffentlicht: (2024)
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2025)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2025)
NRR-Phi: Text-to-State Mapping for Ambiguity Preservation in LLM Inference
von: Saito, Kei
Veröffentlicht: (2026)
von: Saito, Kei
Veröffentlicht: (2026)
Next Token Prediction Is a Dead End for Creativity
von: Olatunji, Ibukun, et al.
Veröffentlicht: (2025)
von: Olatunji, Ibukun, et al.
Veröffentlicht: (2025)
RWKV-7 "Goose" with Expressive Dynamic State Evolution
von: Peng, Bo, et al.
Veröffentlicht: (2025)
von: Peng, Bo, et al.
Veröffentlicht: (2025)
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
von: Dhole, Kaustubh D.
Veröffentlicht: (2026)
von: Dhole, Kaustubh D.
Veröffentlicht: (2026)
ACCORD: Closing the Commonsense Measurability Gap
von: Roewer-Després, François, et al.
Veröffentlicht: (2024)
von: Roewer-Després, François, et al.
Veröffentlicht: (2024)
Enhancing Transformer RNNs with Multiple Temporal Perspectives
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey
von: Li, Caihua, et al.
Veröffentlicht: (2026)
von: Li, Caihua, et al.
Veröffentlicht: (2026)
Perturbation Dose Responses in Recursive LLM Loops: Raw Switching, Stochastic Floors, and Persistent Escape under Append, Replace, and Dialog Updates
von: Kaplanski, Pawel
Veröffentlicht: (2026)
von: Kaplanski, Pawel
Veröffentlicht: (2026)
Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
von: Ackermann, Richard, et al.
Veröffentlicht: (2025)
von: Ackermann, Richard, et al.
Veröffentlicht: (2025)
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
von: Vegner, Ivan, et al.
Veröffentlicht: (2025)
von: Vegner, Ivan, et al.
Veröffentlicht: (2025)
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
von: Baxi, Rahul
Veröffentlicht: (2025)
von: Baxi, Rahul
Veröffentlicht: (2025)
Graph Language Models
von: Plenz, Moritz, et al.
Veröffentlicht: (2024)
von: Plenz, Moritz, et al.
Veröffentlicht: (2024)
AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by gpt-4o-mini
von: Rettberg, Jill Walker, et al.
Veröffentlicht: (2025)
von: Rettberg, Jill Walker, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
von: Teixeira, Tiago, et al.
Veröffentlicht: (2026) -
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
von: Du, Bangde, et al.
Veröffentlicht: (2025) -
ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration
von: Deng, Minghang, et al.
Veröffentlicht: (2025) -
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
von: Williamson, Dane, et al.
Veröffentlicht: (2025) -
LLMs and the Human Condition
von: Wallis, Peter
Veröffentlicht: (2024)