# Current Evaluation Evidence

This directory is the compact evidence index for
[`Current_Implementation_Evaluation_2026-07-30.md`](../../Current_Implementation_Evaluation_2026-07-30.md).
The raw run directories and benchmark binaries live in the companion
`.symcc-research-evidence/current-eval-2026-07-30/` tree. This directory alone
is therefore not a sealed source-and-binary snapshot.

The run directory names retain the early `paired_15s_*` label for provenance.
Repeat indices were not randomized blocks and have no designed pairing
semantics. The current analysis treats the two campaigns as independent
samples. `paired_comparisons.csv` is audit history only and must not be used for
the report's claims.

## Campaigns

All commands were run from `/home/ubuntu/code/symcc` on 2026-07-30 UTC.

Legacy orchestration, 20 repeats per target and mode:

```bash
python3 benchmark/run_benchmark.py \
  --symcc build/symcc --np-list 8 --rounds 20 --timeout 15 \
  --output .symcc-research-evidence/current-eval-2026-07-30/runs/paired_15s_legacy \
  --bin-dir .symcc-research-evidence/current-eval-2026-07-30/bin \
  --skip-build --no-default --hybrid --afl-only --no-serial --no-mpi \
  --hybrid-afl-instances 1 --aflpp-profiles off \
  --public sqlite-sqlite_fuzzer:/home/ubuntu/code/symcc/.symcc-research-evidence/current-eval-2026-07-30/bin/sqlite_fuzzer:/home/ubuntu/code/symcc/benchmark/public/seeds/sqlite/sqlite_fuzzer \
  libarchive-archive_fuzzer:/home/ubuntu/code/symcc/.symcc-research-evidence/current-eval-2026-07-30/bin/archive_fuzzer:/home/ubuntu/code/symcc/benchmark/public/seeds/libarchive/archive_fuzzer
```

Current-default orchestration, 20 repeats per target and mode:

```bash
python3 benchmark/run_benchmark.py \
  --symcc build/symcc --np-list 8 --rounds 20 --timeout 15 \
  --output .symcc-research-evidence/current-eval-2026-07-30/runs/paired_15s_current \
  --bin-dir .symcc-research-evidence/current-eval-2026-07-30/bin \
  --skip-build --no-default --hybrid --afl-only --no-serial --no-mpi \
  --hybrid-afl-instances 0 --aflpp-profiles basic \
  --public sqlite-sqlite_fuzzer:/home/ubuntu/code/symcc/.symcc-research-evidence/current-eval-2026-07-30/bin/sqlite_fuzzer:/home/ubuntu/code/symcc/benchmark/public/seeds/sqlite/sqlite_fuzzer \
  libarchive-archive_fuzzer:/home/ubuntu/code/symcc/.symcc-research-evidence/current-eval-2026-07-30/bin/archive_fuzzer:/home/ubuntu/code/symcc/benchmark/public/seeds/libarchive/archive_fuzzer
```

Current-default descriptive 300 s run:

```bash
python3 benchmark/run_benchmark.py \
  --symcc build/symcc --np-list 8 --rounds 1 --timeout 300 \
  --output .symcc-research-evidence/current-eval-2026-07-30/runs/descriptive_300s_current \
  --bin-dir .symcc-research-evidence/current-eval-2026-07-30/bin \
  --skip-build --no-default --hybrid --afl-only --no-serial --no-mpi \
  --hybrid-afl-instances 0 --aflpp-profiles basic \
  --public sqlite-sqlite_fuzzer:/home/ubuntu/code/symcc/.symcc-research-evidence/current-eval-2026-07-30/bin/sqlite_fuzzer:/home/ubuntu/code/symcc/benchmark/public/seeds/sqlite/sqlite_fuzzer \
  libarchive-archive_fuzzer:/home/ubuntu/code/symcc/.symcc-research-evidence/current-eval-2026-07-30/bin/archive_fuzzer:/home/ubuntu/code/symcc/benchmark/public/seeds/libarchive/archive_fuzzer
```

Statistical analysis:

```bash
python3 benchmark/analyze_current_evaluation.py \
  --group legacy=docs/evidence/current-eval-2026-07-30/paired_15s_legacy.csv \
  --group current=docs/evidence/current-eval-2026-07-30/paired_15s_current.csv \
  --comparison current/hybrid,legacy/hybrid \
  --comparison legacy/hybrid,legacy/afl-only \
  --comparison current/hybrid,current/afl-only \
  --comparison current/afl-only,legacy/afl-only \
  --output-dir /tmp/current-evaluation-analysis
```

The report uses `independent_comparisons.csv`, generated by independent
bootstrap mean differences and a two-sided label-permutation test with Holm
family-wise correction. `statistical_summary.md` records the full method and
all comparisons.

## Binary identities

| Artifact | SHA-256 |
|---|---|
| `build/symcc` | `f196e5920bd96689db4a7224dbc1c92b4f1c1f99701726e39c0d6aab9a460e97` |
| `build/libsymcc.so` | `c74d4252c41e355a05ae498b75c03f8a4be3b432ec468073fb93afe8f71d1cd0` |
| current SQLite target | `0b0c2d16d0ddd10d9e4555ad6a650f91bacad7b9b6579fb3c665468bee8a677a` |
| current libarchive target | `23f8b523e28f8342f766bb133b28f794160447ee598bcbb29f32c30a6959cc7b` |

## Evidence identities

| Artifact | SHA-256 |
|---|---|
| `paired_15s_legacy.csv` | `22a10551ca752756bf0ee6c42c6d9f1ded9596e809d4397cc24160682ce3f9ee` |
| `paired_15s_current.csv` | `3dd501256f5d7fc110b3f885f3fc6174072447feb5c4fc21f36864d422be1275` |
| `descriptive_300s_current.csv` | `c4df7caf8442da543d729f68e4ce61348627307f33b63e399396b953aa02a85e` |
| `independent_comparisons.csv` | `b977fc7c8d4f1c138d11a147d3220f5a293cfb87e4e7ed0f727e025a45c7eb70` |
| `paired_comparisons.csv` (audit history only) | `6fee57589aa46ddf8a6f573a414523c2e14a3e6010e7b0a40b2916b3a0aa3184` |
| `statistical_results.json` | `bb0555b7dddbe263e8d124a8bf0865642eee7efaafd9c95bfef37207c0be92b2` |
| `statistical_summary.md` | `9992286867398a97047f092ba481f39665e002e6bfd45879b72f1e36bf810471` |
| `edge_coverage_distributions.svg` | `ddd472acdc70384ef28ec1a19dda0b27def727d0e9e97d862c869b82a87cb960` |
| `edge_coverage_distributions.png` | `16af6f1e99a66b364cf7c404b7c01566f01c8fd9668b4e9bbdc3c3ce5c56cdb3` |
| `verification.md` | `80ac1b15150ba299fb1b6738337b9a97540d368b1732fae07eef57217eff55c3` |
| crash reproduction log | `33a89afeb0151c10de17baa0f3e8f56fa7597ddf44b5e45062ac787295d6ee19` |
| post-fix compile log | `d3ce9a54d6bdf94ccd31d36fac21bea8173fb5b1feb4689ca2366ed567f9dc93` |

The 15 s campaigns predate the CSV contribution-field patch, so their total
`generated` counts cannot be losslessly split after the fact. The 300 s CSV was
generated after the patch and records `afl_generated`, `symcc_generated`, and
`symcc_interesting`. A subsequent audit found that its `symcc_interesting`
parser selected the first progress log rather than the final cumulative log;
the raw 47/50 values are therefore retained but must not be reported as final.
The parser now prefers `Final stats` and falls back to `.symcc_stats` and the
accepted queue count. In the two AFL-only rows, `generated` is the AFL count;
their newly added `afl_generated` column is zero because the caller-side
assignment was fixed immediately after this campaign. No raw result was
rewritten retroactively.

## Reproducibility boundary

The campaigns record commands, target/runtime binary hashes, raw CSV rows and
analysis code. They predate a sealed randomized-block protocol, and the current
repository contains uncommitted and untracked implementation work. The evidence
therefore supports an **A-level engineering comparison**, not an R-level
confirmatory result. A future confirmatory campaign must seal the complete
source tree and toolchain, randomize/interleave blocks, equalize CPU-seconds,
and preserve every run manifest before execution.
