Tape: A Cellular Automata Benchmark for Evaluating Rule-Shift Generalization in Reinforcement Learning
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Pan, Enze |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
RASP-Tuner: Retrieval-Augmented Soft Prompts for Context-Aware Black-Box Optimization in Non-Stationary Environments
par: Pan, Enze
Publié: (2026)
par: Pan, Enze
Publié: (2026)
Understanding Knowledge Transferability for Transfer Learning: A Survey
par: Wang, Haohua, et autres
Publié: (2025)
par: Wang, Haohua, et autres
Publié: (2025)
A Theoretical Analysis of Soft-Label vs Hard-Label Training in Neural Networks
par: Mandal, Saptarshi, et autres
Publié: (2024)
par: Mandal, Saptarshi, et autres
Publié: (2024)
Dynamic Observation Policies in Observation Cost-Sensitive Reinforcement Learning
par: Bellinger, Colin, et autres
Publié: (2023)
par: Bellinger, Colin, et autres
Publié: (2023)
seqme: a Python library for evaluating biological sequence design
par: Møller-Larsen, Rasmus, et autres
Publié: (2025)
par: Møller-Larsen, Rasmus, et autres
Publié: (2025)
Reducing Instability in Synthetic Data Evaluation with a Super-Metric in MalDataGen
par: da Silva, Anna Luiza Gomes, et autres
Publié: (2025)
par: da Silva, Anna Luiza Gomes, et autres
Publié: (2025)
MH-FSF: A Unified Framework for Overcoming Benchmarking and Reproducibility Limitations in Feature Selection Evaluation
par: Rocha, Vanderson, et autres
Publié: (2025)
par: Rocha, Vanderson, et autres
Publié: (2025)
confopt: A Library for Implementation and Evaluation of Gradient-based One-Shot NAS Methods
par: Jha, Abhash Kumar, et autres
Publié: (2025)
par: Jha, Abhash Kumar, et autres
Publié: (2025)
Learning Through Noise: Why Subliminal Learning Works and When It Fails
par: Brockers, Vincent C., et autres
Publié: (2026)
par: Brockers, Vincent C., et autres
Publié: (2026)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
par: Aksoy, Sinan G., et autres
Publié: (2026)
par: Aksoy, Sinan G., et autres
Publié: (2026)
Fast, close, non-singular and property-preserving approximations of entropic measures
par: Horenko, Illia, et autres
Publié: (2025)
par: Horenko, Illia, et autres
Publié: (2025)
Achieving Distributive Justice in Federated Learning via Uncertainty Quantification
par: Carey, Alycia, et autres
Publié: (2025)
par: Carey, Alycia, et autres
Publié: (2025)
FedUNet: A Lightweight Additive U-Net Module for Federated Learning with Heterogeneous Models
par: Seo, Beomseok, et autres
Publié: (2025)
par: Seo, Beomseok, et autres
Publié: (2025)
Return of the Schema: Building Complete Datasets for Machine Learning and Reasoning on Knowledge Graphs
par: Diliso, Ivan, et autres
Publié: (2026)
par: Diliso, Ivan, et autres
Publié: (2026)
Verifiable evaluations of machine learning models using zkSNARKs
par: South, Tobin, et autres
Publié: (2024)
par: South, Tobin, et autres
Publié: (2024)
Auditable Homomorphic-based Decentralized Collaborative AI with Attribute-based Differential Privacy
par: Yeh, Lo-Yao, et autres
Publié: (2024)
par: Yeh, Lo-Yao, et autres
Publié: (2024)
One Prompt is not Enough: Automated Construction of a Mixture-of-Expert Prompts
par: Wang, Ruochen, et autres
Publié: (2024)
par: Wang, Ruochen, et autres
Publié: (2024)
What is the $\textit{intrinsic}$ dimension of your binary data? -- and how to compute it quickly
par: Hanika, Tom, et autres
Publié: (2024)
par: Hanika, Tom, et autres
Publié: (2024)
Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
par: Turan, Berkant, et autres
Publié: (2025)
par: Turan, Berkant, et autres
Publié: (2025)
Benchmarking PNW Model for MedMNIST to 100% Accuracy
par: Deng, Bo
Publié: (2026)
par: Deng, Bo
Publié: (2026)
Graph Transformers: A Survey
par: Shehzad, Ahsan, et autres
Publié: (2024)
par: Shehzad, Ahsan, et autres
Publié: (2024)
On the Invariants of Softmax Attention
par: Lee, Wonsuk
Publié: (2026)
par: Lee, Wonsuk
Publié: (2026)
ATEX-CF: Attack-Informed Counterfactual Explanations for Graph Neural Networks
par: Zhang, Yu, et autres
Publié: (2026)
par: Zhang, Yu, et autres
Publié: (2026)
Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading
par: Sadjoli, Nicholas, et autres
Publié: (2026)
par: Sadjoli, Nicholas, et autres
Publié: (2026)
The Positivity of the Neural Tangent Kernel
par: Carvalho, Luís, et autres
Publié: (2024)
par: Carvalho, Luís, et autres
Publié: (2024)
ShapG: new feature importance method based on the Shapley value
par: Zhao, Chi, et autres
Publié: (2024)
par: Zhao, Chi, et autres
Publié: (2024)
Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation
par: Song, Sum Kyun, et autres
Publié: (2026)
par: Song, Sum Kyun, et autres
Publié: (2026)
Benchmarking Deception Probes via Black-to-White Performance Boosts
par: Parrack, Avi, et autres
Publié: (2025)
par: Parrack, Avi, et autres
Publié: (2025)
Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark
par: Yu, Zhiqi, et autres
Publié: (2026)
par: Yu, Zhiqi, et autres
Publié: (2026)
Enterprise Resource Planning Using Multi-type Transformers in Ferro-Titanium Industry
par: Yazdanpourmoghadam, Samira, et autres
Publié: (2026)
par: Yazdanpourmoghadam, Samira, et autres
Publié: (2026)
Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks
par: Paim, Kayua Oleques, et autres
Publié: (2025)
par: Paim, Kayua Oleques, et autres
Publié: (2025)
Score Change of Variables
par: Robbins, Stephen
Publié: (2024)
par: Robbins, Stephen
Publié: (2024)
On-Premise SLMs vs. Commercial LLMs: Prompt Engineering and Incident Classification in SOCs and CSIRTs
par: Almeida, Gefté, et autres
Publié: (2025)
par: Almeida, Gefté, et autres
Publié: (2025)
LAMBDA: A Large Model Based Data Agent
par: Sun, Maojun, et autres
Publié: (2024)
par: Sun, Maojun, et autres
Publié: (2024)
A New Similarity Function for Spectral Clustering with Application to Plant Phenotypic Data
par: Ahuja, Kapil, et autres
Publié: (2023)
par: Ahuja, Kapil, et autres
Publié: (2023)
The Selective G-Bispectrum and its Inversion: Applications to G-Invariant Networks
par: Mataigne, Simon, et autres
Publié: (2024)
par: Mataigne, Simon, et autres
Publié: (2024)
Taxonomy to Regulation: A (Geo)Political Taxonomy for AI Risks and Regulatory Measures in the EU AI Act
par: Arda, Sinan
Publié: (2024)
par: Arda, Sinan
Publié: (2024)
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
par: Lim, Soohan, et autres
Publié: (2025)
par: Lim, Soohan, et autres
Publié: (2025)
Representation Integrity in Temporal Graph Learning Methods
par: Kooshafar, Elahe
Publié: (2025)
par: Kooshafar, Elahe
Publié: (2025)
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
par: Shafieinejad, Masoumeh, et autres
Publié: (2026)
par: Shafieinejad, Masoumeh, et autres
Publié: (2026)
Documents similaires
-
RASP-Tuner: Retrieval-Augmented Soft Prompts for Context-Aware Black-Box Optimization in Non-Stationary Environments
par: Pan, Enze
Publié: (2026) -
Understanding Knowledge Transferability for Transfer Learning: A Survey
par: Wang, Haohua, et autres
Publié: (2025) -
A Theoretical Analysis of Soft-Label vs Hard-Label Training in Neural Networks
par: Mandal, Saptarshi, et autres
Publié: (2024) -
Dynamic Observation Policies in Observation Cost-Sensitive Reinforcement Learning
par: Bellinger, Colin, et autres
Publié: (2023) -
seqme: a Python library for evaluating biological sequence design
par: Møller-Larsen, Rasmus, et autres
Publié: (2025)