Software Artifact & Technical Guides¶
- ACM Artifact Review and Badging - The current badging guidelines for judging scientific software artifacts.
- Artifact Evaluation: Tips for Authors - Ten experience-based tips, with justification and examples, for creating software artifacts.
- BenchExec - A framework for reliable benchmarking of non-interactive tools, with built-in resource control and a table generator for visualizing results.
- Benchmarking Crimes - A synopsis of the many ways an experiment design or analysis can go wrong.
- Can you trust your experimental results? - A general framework for validating experimental designs; a technical report developed based on the Evaluate 2011 workshop.
- EAPLS Artifact Badges - The European badging scheme for software artifact evaluation.
- Empirical Evaluation Guidelines - A checklist to evaluate soundness of scientific experiment setup, developed.
- Empirical Standards for Software Engineering research - The official evidence standards for conducting and reporting studies in software engineering; developed.
- Guide for Accelerating Computational Reproducibility in the Social Sciences - A structured guidebook toward assessing and improving computational reproducibility.
- Handbook for Reproduction and Replication Studies - A practical how-to guide for how to carry out a reproduction or a replication study.
- Proof Artifacts: Guidelines for Submission and Reviewing - Proof artifacts are a special category of scientific software and thus have their own presentation standards; the guidelines are maintained.
- Reliable benchmarking: requirements and solutions - Motivations for reliable benchmarking and presentation of BenchExec.
- STABILIZER: statistically sound performance evaluation - Addressing the bias that commonly arises in measurements of effect size, i.e., the magnitude of change in performance.
- Scientific benchmarking of parallel computing systems: twelve ways to tell the masses when reporting performance results - 12 rules of best practices for reporting empirical evaluation results.