Benchmarks / run 2026-09-27-38b7137-gh31
Ten languages, five evaluations, every result published
The Quidra project runs one frozen methodology over ten languages, Quidra included, and publishes the whole result as data, including the evaluations where Quidra currently ranks last. Rankings are descriptive of the frozen sample; the runner computes them mechanically from validated results, and no cross-evaluation total is ever produced.
- Evaluated compiler
- Quidra 0.3.0, benchmark snapshot
38b7137; its compiler source is identical to release v0.3.0 - Run
2026-09-27-38b7137-gh31(2026-09-27)- Inference model
claude-sonnet-5, used only where a judgment is required- Toolchains
- Quidra 0.3.0 · Python 3.12.3 · C++ 18.1.3 · Rust 1.95.0 · Go 1.26.3 · Java 26.0.1 · TypeScript 7.0.2 · Kotlin 2.3.21 · Swift 6.2.3 · Zig 0.16.0 (compiler or runtime versions from the frozen configuration)
Quidra's position in each evaluation
01 / complete
Semantic Compression
How much program meaning a language states explicitly and determinately per unit of syntax, relative to the breadth of capabilities it can express. All ten languages are annotated against one frozen, language-neutral semantic-site matrix over 44 probes. Five quality metrics (density, determinacy, locality, hidden semantic cost, capability efficiency) form a quality score, and the published score is its harmonic mean with capability coverage, so a tiny language cannot score high by having few rules. Annotation is worker judgment checked by a mechanical validator and a blinded comparability audit. Higher is better.
| Rank | Language | Score |
|---|---|---|
| 1 | Zig | 78.48 |
| 2 | Quidra | 77.52 |
| 3 | Swift | 77.34 |
| 4 | Go | 72.08 |
| 5 | Kotlin | 69.24 |
| 6 | Rust | 67.91 |
| 7 | Java | 53.48 |
| 8 | TypeScript | 48.19 |
| 9 | Python | 46.99 |
| 10 | C++ | 20.38 |
Per-requirement scores (6)
| Requirement | Weight | Quidra | Python | C++ | Rust | Go | Java | TypeScript | Kotlin | Swift | Zig |
|---|---|---|---|---|---|---|---|---|---|---|---|
metric.semantic_density | 0.2 | 100.00 | 23.42 | 4.96 | 29.99 | 45.65 | 0.00 | 10.88 | 58.52 | 70.49 | 58.69 |
metric.semantic_determinacy | 0.25 | 100.00 | 64.75 | 0.00 | 70.50 | 41.00 | 70.50 | 80.33 | 64.75 | 80.33 | 45.08 |
metric.semantic_locality | 0.2 | 34.45 | 45.93 | 0.00 | 28.71 | 91.86 | 17.22 | 40.19 | 28.71 | 0.00 | 100.00 |
metric.hidden_semantic_cost | 0.2 | 81.48 | 18.52 | 11.11 | 66.67 | 48.15 | 48.15 | 0.00 | 66.67 | 74.07 | 100.00 |
metric.capability_efficiency | 0.15 | 36.46 | 0.00 | 54.31 | 60.12 | 82.20 | 47.90 | 24.43 | 59.63 | 100.00 | 27.91 |
metric.capability_coverage | — | 81.82 | 77.27 | 98.86 | 98.86 | 90.91 | 90.91 | 82.95 | 90.91 | 97.73 | 94.32 |
02 / complete
LLM Learnability
Whether a model can learn a language's rules from a reference pack and apply them when familiarity is controlled: the language is never named, keywords or vocabulary are transformed by reversible mappings in most conditions, and every program is mapped back and validated by the real toolchain. Six conditions (keyword and vocabulary anonymization, surface perturbation, novel-rule generalization, held-out rule composition, prior-conflict resistance) are scored with the same task metrics as LLM Proficiency and combined with fixed weights. It is separate from LLM Proficiency and is not a proof of zero prior exposure. Higher is better.
| Rank | Language | Score |
|---|---|---|
| 1 | Swift | 100.00 |
| 2 | TypeScript | 99.10 |
| 3 | Rust | 98.40 |
| 4 | Python | 98.02 |
| 5 | Java | 97.80 |
| 6 | Kotlin | 97.60 |
| 7 | C++ | 97.50 |
| 8 | Go | 91.58 |
| 9 | Zig | 81.67 |
| 10 | Quidra | 71.97 |
Per-requirement scores (6)
| Requirement | Weight | Quidra | Python | C++ | Rust | Go | Java | TypeScript | Kotlin | Swift | Zig |
|---|---|---|---|---|---|---|---|---|---|---|---|
condition.i1_keyword_anonymization | 0.2 | 96.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 67.67 |
condition.i2_vocabulary_anonymization | 0.2 | 57.33 | 100.00 | 92.00 | 100.00 | 66.67 | 100.00 | 100.00 | 100.00 | 100.00 | 67.67 |
condition.i3_structural_surface_perturbation | 0.15 | 100.00 | 100.00 | 100.00 | 96.00 | 95.00 | 92.00 | 100.00 | 96.00 | 100.00 | 92.00 |
condition.i4_novel_rule_generalization | 0.2 | 100.00 | 100.00 | 100.00 | 95.00 | 95.00 | 95.00 | 100.00 | 95.00 | 100.00 | 92.00 |
condition.i5_held_out_rule_composition | 0.15 | 34.00 | 91.50 | 96.00 | 100.00 | 100.00 | 100.00 | 96.00 | 96.00 | 100.00 | 88.00 |
condition.i6_prior_conflict_resistance | 0.1 | 12.00 | 93.00 | 97.00 | 100.00 | 100.00 | 100.00 | 97.00 | 98.00 | 100.00 | 92.00 |
03 / complete
Language Quality
The intrinsic quality of the language and its implementation on comparable workloads written idiomatically in each language, excluding maturity and network effects, which belong to Ecosystem. Twenty-five metrics in four weighted categories: performance and resource metrics are mechanical measurements of the same programs; safety and robustness metrics are computed from frozen adversarial case sets that record how early a defect is detected; language and development metrics are worker judgments on frozen rubrics. Scores are 0-100, higher is better.
| Rank | Language | Score |
|---|---|---|
| 1 | Rust | 72.24 |
| 2 | Go | 67.22 |
| 3 | Zig | 66.76 |
| 4 | Quidra | 66.13 |
| 5 | Python | 64.40 |
| 6 | Swift | 63.64 |
| 7 | Java | 58.77 |
| 8 | Kotlin | 56.95 |
| 9 | C++ | 51.32 |
| 10 | TypeScript | 46.72 |
Per-requirement scores (25)
| Requirement | Weight | Quidra | Python | C++ | Rust | Go | Java | TypeScript | Kotlin | Swift | Zig |
|---|---|---|---|---|---|---|---|---|---|---|---|
metric.native_execution_performance | — | 63.02 | 3.77 | 87.21 | 84.00 | 76.26 | 62.82 | 42.16 | 61.43 | 71.08 | 97.01 |
metric.long_running_performance | — | 61.81 | 3.68 | 85.85 | 82.04 | 74.34 | 74.55 | 41.44 | 70.54 | 69.53 | 95.32 |
metric.compile_build_performance | — | 0.60 | 100.00 | 0.07 | 0.56 | 0.77 | 0.16 | 0.30 | 0.02 | 0.14 | 0.01 |
metric.startup_latency | — | 64.10 | 11.07 | 63.08 | 92.44 | 64.76 | 4.05 | 4.95 | 2.49 | 33.46 | 100.00 |
metric.memory_efficiency | — | 75.17 | 64.84 | 92.07 | 95.93 | 83.55 | 21.87 | 21.87 | 21.35 | 75.94 | 97.81 |
metric.source_code_size | — | 92.32 | 95.13 | 63.24 | 79.87 | 91.21 | 75.05 | 60.87 | 84.04 | 71.11 | 61.20 |
metric.binary_artifact_size | — | 1.70 | 100.00 | 17.09 | 0.09 | 0.17 | 52.78 | 55.90 | 0.07 | 11.36 | 0.11 |
metric.runtime_overhead | — | 100.00 | 99.72 | 99.65 | 99.68 | 99.68 | 26.19 | 22.10 | 24.27 | 99.72 | 99.68 |
metric.code_efficiency_conciseness | — | 90.00 | 90.00 | 50.00 | 75.00 | 60.00 | 50.00 | 90.00 | 95.00 | 95.00 | 65.00 |
metric.readability | — | 95.00 | 85.00 | 50.00 | 80.00 | 95.00 | 85.00 | 75.00 | 90.00 | 90.00 | 75.00 |
metric.functionality_expressiveness | — | 85.00 | 90.00 | 90.00 | 100.00 | 75.00 | 85.00 | 90.00 | 95.00 | 95.00 | 80.00 |
metric.diagnostics | — | 100.00 | 65.00 | 60.00 | 100.00 | 65.00 | 75.00 | 95.00 | 85.00 | 95.00 | 95.00 |
metric.dependency_simplicity | — | 100.00 | 60.00 | 35.00 | 100.00 | 100.00 | 75.00 | 65.00 | 65.00 | 95.00 | 95.00 |
metric.portability_design_platform_neutrality | — | 90.00 | 90.00 | 50.00 | 90.00 | 100.00 | 95.00 | 80.00 | 100.00 | 80.00 | 100.00 |
metric.ffi_interoperability_design | — | 70.00 | 70.00 | 80.00 | 100.00 | 65.00 | 85.00 | 30.00 | 85.00 | 95.00 | 95.00 |
metric.concurrency | — | 55.00 | 60.00 | 60.00 | 90.00 | 70.00 | 85.00 | 75.00 | 85.00 | 100.00 | 65.00 |
metric.type_safety | — | 74.58 | 29.17 | 19.58 | 51.67 | 62.50 | 45.83 | 33.33 | 56.25 | 67.08 | 67.08 |
metric.memory_safety | — | 44.44 | 66.67 | 11.11 | 77.78 | 88.89 | 100.00 | 55.56 | 100.00 | 44.44 | 33.33 |
metric.runtime_safety | — | 41.67 | 68.75 | 3.33 | 39.02 | 51.16 | 57.69 | 22.22 | 45.83 | 13.95 | 9.76 |
metric.boundary_value_safety | — | 54.55 | 65.91 | 7.27 | 51.93 | 60.45 | 42.27 | 22.27 | 37.73 | 47.27 | 53.75 |
metric.adversarial_input_robustness | — | 50.71 | 79.29 | 48.57 | 69.29 | 64.29 | 85.71 | 59.29 | 75.00 | 51.43 | 42.86 |
metric.early_error_detection | — | 59.81 | 57.88 | 25.38 | 64.47 | 64.04 | 65.00 | 44.62 | 61.15 | 55.96 | 49.86 |
metric.debuggability | — | 66.67 | 68.59 | 13.46 | 54.49 | 45.83 | 61.54 | 38.46 | 52.24 | 32.05 | 27.24 |
metric.silent_bug_resistance | — | 91.18 | 70.59 | 29.41 | 75.00 | 72.06 | 70.59 | 52.94 | 64.71 | 88.24 | 63.24 |
metric.implementation_robustness | — | 100.00 | 100.00 | 100.00 | 97.06 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 97.06 |
04 / complete
Ecosystem
The infrastructure, maturity and real-world evidence around a language: libraries, package activity, third-party tools, production adoption, community knowledge, tooling, editor and debugger support, documentation, installation, release maturity and platform coverage. It is kept separate from Language Quality so a young language's maturity gap stays visible. Fifteen metrics in two equally weighted categories; for each, an LLM worker gathers evidence by web search and assigns rubric levels that the runner converts to a 0-100 score. Absent evidence scores low rather than N/A. Higher is better.
| Rank | Language | Score |
|---|---|---|
| 1 | Java | 99.25 |
| 2 | TypeScript | 98.50 |
| 3 | Python | 98.25 |
| 4 | Kotlin | 97.75 |
| 5 | Go | 97.50 |
| 5 | Rust | 97.50 |
| 7 | C++ | 96.25 |
| 8 | Swift | 91.00 |
| 9 | Zig | 66.50 |
| 10 | Quidra | 32.50 |
Per-requirement scores (15)
| Requirement | Weight | Quidra | Python | C++ | Rust | Go | Java | TypeScript | Kotlin | Swift | Zig |
|---|---|---|---|---|---|---|---|---|---|---|---|
metric.third_party_library_availability_domain_coverage | — | 0.00 | 100.00 | 100.00 | 95.00 | 90.00 | 100.00 | 100.00 | 95.00 | 80.00 | 45.00 |
metric.package_ecosystem_activity_maintenance | — | 30.00 | 100.00 | 90.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 85.00 | 50.00 |
metric.third_party_tool_availability | — | 0.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 70.00 |
metric.production_adoption_deployment_evidence | — | 0.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 95.00 | 55.00 |
metric.community_public_knowledge_availability | — | 15.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 95.00 | 90.00 | 60.00 |
metric.core_tooling_availability_quality | — | 75.00 | 100.00 | 95.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 |
metric.package_dependency_management_quality | — | 60.00 | 95.00 | 80.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 70.00 |
metric.ide_editor_support_quality | — | 70.00 | 100.00 | 100.00 | 95.00 | 100.00 | 100.00 | 100.00 | 100.00 | 95.00 | 70.00 |
metric.debugger_profiler_support | — | 45.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 100.00 | 95.00 | 95.00 | 80.00 |
metric.build_test_integration | — | 60.00 | 100.00 | 95.00 | 95.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 85.00 |
metric.documentation_quality | — | 65.00 | 100.00 | 85.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 95.00 | 75.00 |
metric.installation_distribution_experience | — | 60.00 | 100.00 | 90.00 | 100.00 | 100.00 | 100.00 | 100.00 | 95.00 | 95.00 | 80.00 |
metric.toolchain_stability_release_maturity | — | 50.00 | 100.00 | 100.00 | 95.00 | 100.00 | 100.00 | 95.00 | 100.00 | 100.00 | 60.00 |
metric.implemented_platform_coverage | — | 40.00 | 70.00 | 100.00 | 95.00 | 85.00 | 85.00 | 85.00 | 90.00 | 85.00 | 85.00 |
metric.external_integration_coverage | — | 35.00 | 100.00 | 100.00 | 90.00 | 85.00 | 100.00 | 90.00 | 95.00 | 85.00 | 75.00 |
05 / complete
LLM Proficiency
How useful a language is to a present-day LLM under its real name, syntax, standard library and toolchain; pretraining familiarity is deliberately included. One fixed model implements three frozen workloads under two scenarios with up to three repair turns, and correctness is decided by hidden-input oracles. Eighteen weighted metrics, including Correct@1, silent-bug resistance, test pass rate and repair success, combine into a 0-100 score; sixteen are computed by the runner from trial evidence and two are worker-judged. Higher is better.
| Rank | Language | Score |
|---|---|---|
| 1 | Python | 88.68 |
| 2 | Rust | 85.28 |
| 3 | Go | 82.57 |
| 4 | C++ | 76.48 |
| 5 | Kotlin | 74.01 |
| 6 | Java | 72.02 |
| 7 | Swift | 64.78 |
| 8 | TypeScript | 30.99 |
| 9 | Zig | 5.24 |
| 10 | Quidra | 4.91 |
Per-requirement scores (18)
| Requirement | Weight | Quidra | Python | C++ | Rust | Go | Java | TypeScript | Kotlin | Swift | Zig |
|---|---|---|---|---|---|---|---|---|---|---|---|
metric.generation_success_rate | 0.04 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 |
metric.compile_parse_success_rate | 0.04 | 0.00 | 88.89 | 66.67 | 77.78 | 72.22 | 94.44 | 0.00 | 72.22 | 72.22 | 0.00 |
metric.correct_at_1 | 0.12 | 0.00 | 88.89 | 66.67 | 77.78 | 72.22 | 83.33 | 0.00 | 72.22 | 50.00 | 0.00 |
metric.correct_at_n | 0.05 | 0.00 | 100.00 | 100.00 | 100.00 | 100.00 | 88.89 | 61.11 | 100.00 | 100.00 | 0.00 |
metric.test_pass_rate | 0.1 | 0.00 | 86.96 | 60.00 | 82.19 | 80.00 | 64.29 | 20.67 | 65.22 | 55.05 | 0.00 |
metric.repair_success_rate | 0.07 | 0.00 | 100.00 | 100.00 | 100.00 | 100.00 | 33.33 | 61.11 | 100.00 | 100.00 | 0.00 |
metric.repair_efficiency | 0.05 | 0.00 | 95.83 | 83.33 | 94.44 | 93.06 | 86.11 | 40.28 | 86.11 | 79.17 | 0.00 |
metric.diagnosis_efficiency | 0.05 | 0.00 | 50.00 | 0.00 | 100.00 | 100.00 | 0.00 | 44.44 | 0.00 | 33.33 | 0.00 |
metric.silent_bug_resistance | 0.12 | 0.00 | 100.00 | 100.00 | 100.00 | 100.00 | 88.24 | 0.00 | 100.00 | 92.31 | 0.00 |
metric.syntax_hallucination_resistance | 0.07 | 0.00 | 100.00 | 90.00 | 77.80 | 72.20 | 94.40 | 100.00 | 72.20 | 50.00 | 0.00 |
metric.specification_compliance | 0.07 | 0.00 | 100.00 | 100.00 | 90.00 | 72.20 | 88.90 | 61.11 | 100.00 | 66.67 | 0.00 |
metric.prompt_robustness | 0.05 | 0.00 | 66.67 | 50.00 | 66.67 | 66.67 | 83.33 | 0.00 | 66.67 | 33.33 | 0.00 |
metric.unseen_case_generalization | 0.06 | 0.00 | 90.48 | 66.67 | 78.57 | 76.19 | 85.71 | 0.00 | 73.81 | 52.38 | 0.00 |
metric.source_token_efficiency | 0.03 | 24.85 | 90.38 | 81.38 | 46.28 | 67.63 | 55.68 | 40.86 | 59.43 | 64.76 | 33.66 |
metric.total_token_efficiency | 0.03 | 5.37 | 90.16 | 69.08 | 57.91 | 70.23 | 67.38 | 18.51 | 58.15 | 49.10 | 7.60 |
metric.generated_code_performance | 0.03 | 0.00 | 9.74 | 76.92 | 95.16 | 67.59 | 1.84 | 3.34 | 1.81 | 20.88 | 0.00 |
metric.generated_code_memory_efficiency | 0.01 | 0.00 | 99.75 | 99.70 | 99.77 | 99.73 | 20.98 | 12.97 | 22.68 | 62.74 | 0.00 |
metric.generated_code_compile_performance | 0.01 | 0.00 | 100.00 | 2.62 | 12.13 | 18.97 | 4.66 | 6.55 | 0.60 | 3.84 | 0.00 |
How to read these numbers
- Descriptive ranking of the frozen Primary sample; it does not claim statistical superiority beyond the measured sample. Tie policy: competition rank on published score.
- Each evaluation stands on its own. A ranking is published only when all ten languages complete that evaluation and it passes its scientific and integrity gates.
- Language models are used only where judgment is required; planning, retries, measurements, normalization, aggregation and ranking are owned by the deterministic runner.
- Machine-readable result files are authoritative. This page is a derived view of
rankings.json; the per-worker evidence and reasoning are committed beside it.