Skip to content

[Leaderboard] Qwen3.8-Flash-Next (DAB scaffold, hints): 70.87% Pass@1 - #99

Open
qiboyu301-crypto wants to merge 1 commit into
ucbepic:mainfrom
qiboyu301-crypto:codex/qwen38flash-leaderboard
Open

[Leaderboard] Qwen3.8-Flash-Next (DAB scaffold, hints): 70.87% Pass@1#99
qiboyu301-crypto wants to merge 1 commit into
ucbepic:mainfrom
qiboyu301-crypto:codex/qwen38flash-leaderboard

Conversation

@qiboyu301-crypto

Copy link
Copy Markdown

This submits Qwen3.8-Flash-Next running the DAB DataAgent scaffold for leaderboard review.

  • Backbone: Qwen/Qwen3.8-Flash-Next, self-hosted with vLLM on 8 NVIDIA L40S GPUs.
  • Hints: Yes. No added task-specific prompt; the Qwen adapter uses the scaffold's existing tool-result access instruction.
  • Coverage: 54 queries × 5 runs = 270 trials across 12 datasets.
  • Dataset-weighted Pass@1: 0.7086721612 (70.87%). Unweighted trial average: 201/270 = 74.44%.
  • One fixed configuration, no trial selection or replacement runs. Five context-limit failures are retained as empty answers and non-passes.

The submission directory is leaderboard_submissions/qwen38flash_parallel_v1/. It contains the 270 verbatim answers, per-trial scores, failure records, run configuration, deployment metadata, archive checksum, reproduction instructions, and an independent verification script.

The complete execution traces, generated code/intermediates, batch/session records, and actual harness source overlay are available in the trace release. The archive includes all 270 final-agent, LLM-call and tool-call logs. Its manifest records original and published hashes; credentials and text-file deployment locations are redacted, with answer strings unchanged.

Validation

  • Official questions, description/hint files, validators and 54 ground-truth files match checked revision 881bab89f69ac4f150e2586357ede8c5b95719d5.
  • All 270 keys are present exactly once; exported answers match the original final records and corresponding return-answer tool calls.
  • The archive checksum and every published artifact hash are checked by the supplied verifier.
  • The final verification run checked all 3982 archived files and all 270 answer/trace pairs, then independently re-scored with the official validators: 201 passes and dataset-weighted Pass@1 0.7086721611721613.
  • Full run configuration and environment are recorded. The exact upstream weight revision was not captured; this is explicitly disclosed, with available model metadata hashes and serving-image identification supplied instead.

Failures and data-access disclosure

agnews/query4/run1–4 and music_brainz_20k/query3/run2 exhausted the 262144-token context budget while requesting 131072 output tokens. Their complete traces are retained and all five score zero.

The audit scanned 5603 raw tool calls and inspected filesystem-related matches. stockmarket/query2/run0 and run4 enumerated repository filenames, including gold/grader filenames; run0 also counted records from its database configuration. Run2 enumerated directories and read /proc/self/cmdline. We did not identify reading gold/validator contents or obtaining an external answer, but DuckDB host filesystem access was not structurally restricted. All three of these stockmarket trials passed the validators. The calls are retained in the traces and described in the README for rubric review; the audit is not claimed to be exhaustive semantic verification.

Please review the disclosed traces under the submission rubric and independently validate the reported score.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants