[Leaderboard] Stela - deepseek v4.1 flash max - 81.14% Pass@1 - #103
Open
ay27 wants to merge 1 commit into
Open
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stela — Leaderboard Submission
deepseek-flash(recorded endpoint model identifier, DeepSeek-V4.1-Flash)maxdb_description_withhint.txt)Architecture
Stela is a local-first data analysis application. This submission uses its Agent harness with SQL and MongoDB query tools, a per-trial Python workspace, and model-backed semantic operations for text analysis. The agent explores schemas and data, performs joins, transformations, aggregation and ranking, and uses intermediate results to guide subsequent tool calls.
Analysis evidence contracts help track the intended result shape, coverage and supporting evidence. Strategy review supports reassessment during a query. Semantic operations use configured record, request and token budgets. These capabilities were enabled for this evaluation; they are available to the agent rather than guaranteed to run on every trial.
Evaluation
The submission contains the complete 54-query leaderboard suite with five trials per query. All 270 outcomes, including failed outcomes, are retained; this is a full run rather than a selected repair run.
Pass@1 is computed by averaging success across trials for each query, averaging the query scores within each dataset, and then averaging equally across the 12 datasets. This yields 81.14%. The separate trial-weighted success rate is 227/270 (84.07%).
Run configuration
953e4574a9949f79c61bda12914493648fa97505881bab89f69ac4f150e2586357ede8c5b95719d53; MongoDB concurrency:2; Python concurrency:21200000ms; task timeout:3600000ms1000records,200requests,200000tokensConcurrent trials used separate Python workspaces while sharing benchmark database fixtures. Original message timestamps are preserved in the traces.
The model name above is the endpoint identifier recorded by the runner; an exact backend model revision was not independently verified. The deployed source included one additional retained file,
src/components/ai/prompt-input-keyboard.ts, beyond the base commit. The exact deployed source snapshot is retained separately for provenance review.Execution traces
stela-leaderboard-submission.zipcontains:submission.json: the submitted answers.traces.jsonl: 270 records, one per trial, keyed by dataset, query and run.README.md: trace format, verification details and recording limitations.Each trace retains the original ordered application conversation: user messages, assistant reasoning and text, tool calls with their IDs and arguments, tool results and errors, and the exact submitted final answer. Failed tool calls and subsequent corrections are preserved.
All 270 final answers and 10,781 message contents were checked against the original records. All 6,393 tool calls have matching tool results. Internal scheduler logs, grading results and transport metadata are excluded from the reviewer archive.
Recording scope
These are application conversation traces. Per-request system prompts and tool schemas, and complete raw conversations for semantic subcalls, were not captured. Semantic tool inputs and returned results remain visible in the parent conversation. Original archives and the deployed source snapshot are retained separately for provenance review. Please let us know if additional recording is required for verification.
Submission artifacts
leaderboard_submissions/stela.jsonstela-leaderboard-submission.zipbc1980b6121acf577004819597ceb5027eb96de9e7120b04d60741011d4382cestela-leaderboard-submission.zip