Orivel Orivel
Open menu

Coding

Compare implementation quality, correctness, and practical coding ability.

In this genre, the main abilities being tested are Correctness, Completeness, Code Quality.

Unlike system design, this genre focuses more on whether the answer actually works at the code level than on high-level architecture trade-offs.

A high score here does not guarantee strong product judgment, broad architectural thinking, or clear teaching-oriented explanations.

Strong models here are useful for

implementation, debugging, refactoring, and hands-on programming support.

This genre alone cannot tell you

whether the model is best for architecture review, stakeholder writing, or open-ended ideation.

Data analysis

Coding: GPT-5 mini takes rank 1, but the highest average sits at rank 4

26 scored answers Coding Updated 2026/7/1
1
Claude Fable 5

Anthropic

91
Avg. score
100%
Win Rate
1× 1st place 1 samples
2
GPT-5 mini

OpenAI

82
Avg. score
100%
Win Rate
5× 1st place 5 samples
3
Claude Opus 4.8

Anthropic

81
Avg. score
100%
Win Rate
2× 1st place 2 samples

Average score by model

1 Claude Fable 5
9.06
2 GPT-5 mini
8.22
3 Claude Opus 4.8
8.07
4 GPT-5.5
8.90
5 Claude Sonnet 4.6
7.70
6 Gemini 2.5 Pro
7.95
7 Gemini 2.5 Flash-Lite
7.17
8 Gemini 2.5 Flash
6.84

What we weighted

Correctness 35% Completeness 20% Code Quality 20% Practical Value 15% Instruction Following 10%

Across 37 scored coding answers, the ranking is led by GPT-5 mini: 8.22 average over 5 samples, 5 first-place finishes and a 100% win rate. That makes it both the top-ranked model and one of the best-evidenced here, a clean sweep of its matchups at a light-tier cost. Right behind it, Claude Opus 4.8 ranks 2 with 8.07 over just 2 samples and a perfect 100% record, so read it as a strong but still provisional signal.

Average score and rank order diverge sharply, because win rate (head-to-head firsts) drives the ranking more than the raw average. GPT-5.5 posts the highest average in the genre at 8.9, yet ranks only 4, because over 2 samples it won just 1 for a 50% win rate. GPT-5.4, by contrast, carries the largest body of evidence, 8 samples, with an 8.41 average, 6 firsts and a 75% win rate at rank 3. The leader's edge over rank 2 is only 0.15 points, so the top of the table is tight.

The clearest case of a good average buried by weak head-to-head is Gemini 2.5 Pro: 7.95 average, better than mid-table, but rank 6 on a 0% win rate over 4 samples. Claude Sonnet 4.6 leads the middle at 7.7 (50% over 4 samples, rank 5), trailing the GPT-5 group by roughly 0.5 to 1.2 points. The lighter, faster tiers sit lower: Gemini 2.5 Flash-Lite (7.17), Gemini 2.5 Flash (6.84) and Claude Haiku 4.5 (6.48) trail the leader by 1.0 to 1.7 points. With Correctness weighted highest at 35, ahead of Completeness and Code Quality at 20 each, those gaps point to weaker correctness on the harder tasks rather than just style.

The single biggest caveat is sample size. Claude Opus 4.8 and GPT-5.5 rest on 2 samples each, and most models sit on 3 to 8, so averages can swing on a handful of prompts. The 1.74-point spread from top to bottom is real, but the fine ordering inside the 8-point cluster (GPT-5.5, GPT-5.4, GPT-5 mini, Claude Opus 4.8) should be read as provisional. These are condition-dependent measurements, not a verdict on which model is best at coding overall.

Bottom line

For coding you can rely on today, GPT-5 mini is the most defensible pick: rank 1 on a 100% win rate over 5 samples at light-tier cost. GPT-5.4 is the most thoroughly evidenced higher-end option (8.41 over 8 samples), while GPT-5.5's genre-best 8.9 average and Claude Opus 4.8's rank 2 both rest on 2 samples, so treat them as promising but unproven.

This analysis is derived from Orivel's measured benchmark scores for this genre and is updated periodically. Scores are condition-dependent measurements, not absolute truth.

Top Models in This Genre

This ranking is ordered by average score within this genre only.

Latest Updated: Jun 29, 2026 09:44

#1
Claude Fable 5 Anthropic

Win Rate

100%

Average Score

91
#2
GPT-5 mini OpenAI

Win Rate

100%

Average Score

82
#3
Claude Opus 4.8 Anthropic

Win Rate

100%

Average Score

81
#4
GPT-5.5 OpenAI

Win Rate

50%

Average Score

89
#5
Claude Sonnet 4.6 Anthropic

Win Rate

50%

Average Score

77
#6
Gemini 2.5 Pro Google

Win Rate

0%

Average Score

80
#7
Gemini 2.5 Flash-Lite Google

Win Rate

0%

Average Score

72
#8
Gemini 2.5 Flash Google

Win Rate

0%

Average Score

68

What Is Evaluated in Coding

Scoring criteria and weight used for this genre ranking.

Correctness

35.0%

This criterion is included to check Correctness in the answer. It carries heavier weight because this part strongly shapes the overall result in this genre.

Completeness

20.0%

This criterion is included to check Completeness in the answer. It has meaningful weight because it affects quality in a visible way, even if it is not the only thing that matters.

Code Quality

20.0%

This criterion is included to check Code Quality in the answer. It has meaningful weight because it affects quality in a visible way, even if it is not the only thing that matters.

Practical Value

15.0%

This criterion is included to check Practical Value in the answer. It is weighted more lightly because it supports the main goal rather than defining the genre by itself.

Instruction Following

10.0%

This criterion is included to check Instruction Following in the answer. It is weighted more lightly because it supports the main goal rather than defining the genre by itself.

Recent tasks

Coding

Anthropic Claude Opus 4.8 VS Google Gemini 2.5 Flash

Implement a Deterministic Limit Order Book Simulator

Write a single-file Python 3.11 solution implementing the function process_events(events: list[dict]) -> dict. Do not use external packages. The function must simulate a small exchange limit order book for one instrument. It receives a list of event dictionaries in input order and returns a dictionary with exactly these keys: trades, rejected, book. Event types: 1. New order event: Required fields: type="new", id, side, order_type, qty. side is "buy" or "sell". order_type is "limit" or "market". qty is a positive integer. A limit order also requires price, a positive integer number of cents. Optional field tif is time-in-force: "GTC", "IOC", or "FOK". If absent, use "GTC" for limit orders and "IOC" for market orders. Market orders may not have tif="GTC" and may not rest on the book. 2. Cancel event: Required fields: type="cancel", id. It cancels the remaining quantity of a currently resting order with that id. Matching rules: - The book has bids and asks. Resting buy limit orders are bids; resting sell limit orders are asks. - Price-time priority is mandatory: best price first; for the same price, earlier accepted resting order first. - A buy order matches resting asks while it can cross: market buy crosses any ask; limit buy crosses asks with ask price <= buy limit price. - A sell order matches resting bids while it can cross: market sell crosses any bid; limit sell crosses bids with bid price >= sell limit price. - Each trade quantity is min(incoming remaining quantity, resting remaining quantity). - Trade price is always the resting maker order's limit price, never the incoming order's price. - A trade record must be appended immediately when it happens with exactly these keys: buy_id, sell_id, price, qty, taker_id, maker_id. - Partially filled resting orders keep their original priority with the remaining quantity. Fully filled orders leave the book. Time-in-force behavior: - GTC limit orders rest any unfilled remainder on the book. - IOC orders execute as much as possible immediately, then cancel any remainder. - FOK orders must be completely fillable immediately according to the current book and crossing rules. If not completely fillable, they produce no trades and do not change the book. If completely fillable, execute normally. FOK orders never rest. Validation and rejection rules: - If an event is malformed, reject it without changing the book. Append a rejection record to rejected with keys input_index, event, reason. The reason may be a short human-readable string. - Reject a new order if its id is already used by any previously accepted new order, even if that earlier order has since filled or been canceled. - Reject cancel events for unknown ids or ids that are no longer resting. - Reject non-integer, zero, or negative qty and price values. In Python, bool must not be accepted as an integer for these fields. - Ignore extra fields on otherwise valid events. Return format: - trades: list of trade records in execution order. - rejected: list of rejection records in input order. - book: a dictionary with keys bids and asks. - book["bids"] must list all resting bids sorted by descending price, then original resting time, each as {"id": id, "price": price, "qty": remaining_qty}. - book["asks"] must list all resting asks sorted by ascending price, then original resting time, each as {"id": id, "price": price, "qty": remaining_qty}. Your answer should be complete executable Python code defining process_events. You may include helper classes/functions and a small self-test section guarded by if __name__ == "__main__":, but the core function must not read from stdin or write to stdout.

111
Jun 29, 2026 09:44

Coding

Anthropic Claude Opus 4.8 VS Google Gemini 2.5 Pro

Implement Atomic JSON Patch Application in Python

Write a Python 3.11 implementation of a function named apply_json_patch(document, patch) that applies a JSON Patch-style sequence of operations to a JSON-compatible value and returns the patched value. The input document may be any combination of dict, list, str, int, float, bool, and None. The patch is a list of operation dicts. The implementation must not mutate the original document or any nested object reachable from it. If any operation is invalid, the function must raise a custom exception class named JsonPatchError and leave the original document unchanged. Supported operations are add, remove, replace, move, copy, and test. Use JSON Pointer paths with slash-separated tokens, where the empty string identifies the whole document, tokens decode ~1 as / and ~0 as ~, and any other use of ~ is invalid. For objects, a path token is a key. For arrays, a path token must be a non-negative integer without leading zeros except the single token 0; for add only, the final token may be - to append. The add operation inserts into arrays at an index from 0 through len(array), appends for -, sets an object key, or replaces the whole document at path empty. The remove operation requires the target to exist and deletes it. The replace operation requires the target to exist and replaces it. The move operation requires from and path, removes the value at from and adds it at path, and must reject moving a value into one of its own descendants. The copy operation requires from and path and deep-copies the source value to the target. The test operation requires value and succeeds only if the current target is deeply equal to value, including normal Python equality for numbers and exact equality for strings, booleans, and None. Each operation dict must contain exactly the fields required for that operation plus the op field; unknown fields or missing fields are errors. The function should be deterministic, reasonably efficient, and rely only on the Python standard library. Include any helper functions or classes needed. Do not write a command-line program or use external packages.

164
Jun 15, 2026 09:43

Coding

Anthropic Claude Fable 5 VS OpenAI GPT-5.5

Implement a Dependency-Based Task Scheduler in Python

Write a Python function or class that schedules a list of tasks based on their dependencies. The scheduler should determine the order in which tasks can be executed, grouping tasks that can run in parallel. The input will be a list of dictionaries, where each dictionary represents a task with the following keys: - `id`: A unique string identifier for the task. - `name`: A string name for the task. - `dependencies`: A list of string IDs of tasks that must be completed before this task can start. Your implementation should: 1. Take the list of task dictionaries as input. 2. Return a valid execution plan as a list of lists. Each inner list represents a 'batch' of tasks that can be executed concurrently. The order of batches represents the sequential execution order. The order of task IDs within a batch does not matter. 3. Detect and handle circular dependencies. If a cycle is found, it should raise a `ValueError` with a descriptive message. 4. Detect and handle cases where a dependency ID does not correspond to any existing task. This should also raise a `ValueError`.

163
Jun 12, 2026 09:39

Coding

OpenAI GPT-5.5 VS Google Gemini 2.5 Flash

Rate Limiter with Sliding Window and Burst Allowance

Design and implement a thread-safe rate limiter in a language of your choice (Python, Go, Java, TypeScript, or Rust) that supports the following requirements: 1. **API surface**: Expose at least these operations: - `allow(client_id: str, cost: int = 1) -> bool` — returns whether the request is permitted right now. - `retry_after(client_id: str) -> float` — returns seconds until at least 1 unit of capacity is available (0 if currently allowed). - A constructor that accepts per-client configuration: `rate` (units per second), `burst` (max units stored), and an optional `window_seconds` for sliding-window accounting. 2. **Algorithm**: Implement a hybrid that combines a **token bucket** (for burst tolerance) with a **sliding-window log or counter** (to bound the total requests permitted within `window_seconds`, preventing sustained abuse that a pure token bucket would allow after refills). A request is permitted only if both checks pass. Justify your data-structure choice for the sliding window (exact log vs. weighted two-bucket approximation) and discuss memory/accuracy tradeoffs in a short comment block or accompanying note. 3. **Concurrency**: The limiter will be hit by many threads/goroutines concurrently for the same and different `client_id`s. Avoid a single global lock becoming a bottleneck (e.g., per-client locks or lock striping). Document why your approach is correct under concurrent `allow` calls (no double-spend of tokens, no lost updates). 4. **Time source**: Make the clock injectable so tests are deterministic. Use a monotonic clock by default. 5. **Edge cases to handle explicitly**: - `cost` larger than `burst` (must reject, never block forever). - Clock going backwards or large pauses (e.g., suspended VM): clamp rather than crash, and don't grant unbounded tokens. - First-ever request for a new client (lazy initialization). - Stale client cleanup (memory must not grow unbounded if clients stop calling). - Fractional tokens / sub-millisecond timing. 6. **Tests**: Provide at least 6 unit tests using the injectable clock that cover: basic allow/deny, burst draining and refill, sliding-window cap independent of bucket refill, `cost > burst`, concurrent contention on one client (deterministic property: total permitted in T seconds ≤ rate*T + burst), and stale-client eviction. 7. **Complexity**: State the amortized time complexity of `allow` and the memory complexity per client. Deliver: complete runnable code (single file is fine, but you may split files if you label them clearly), the tests, and a brief design note (max ~250 words) explaining your choices and the precise semantics when the two algorithms disagree.

300
May 12, 2026 09:45

Coding

Anthropic Claude Opus 4.7 VS OpenAI GPT-5.4

Markdown Subset to HTML Converter

Write a Python function `markdown_to_html(markdown_text: str) -> str` that converts a string containing a specific subset of Markdown into its corresponding HTML representation. The function must support the following features: **Block Elements:** 1. **Headers:** Lines starting with `# ` to `###### ` should be converted to `<h1>` to `<h6>` tags. 2. **Unordered Lists:** Lines starting with `- ` should be converted to `<ul>` and `<li>` tags. Nested lists, indented by two spaces per level, must be supported. A list is terminated by a blank line or a different block element. 3. **Code Blocks:** Content enclosed between lines of triple backticks (```) should be converted to `<pre><code>...</code></pre>`. The language specifier on the opening backticks (e.g., ```python) should be ignored. No other Markdown processing should occur inside a code block. 4. **Paragraphs:** Any other text should be wrapped in `<p>` tags. Consecutive lines of text belong to the same paragraph. Paragraphs are separated by one or more blank lines. **Inline Elements:** 1. **Bold & Italic:** `***text***` should be converted to `<strong><em>text</em></strong>`. 2. **Bold:** `**text**` should be converted to `<strong>text</strong>`. 3. **Italic:** `*text*` should be converted to `<em>text</em>`. **Rules and Constraints:** - Inline elements can be nested within headers and list items. - The parser should be robust to malformed or tricky inputs, such as unclosed inline tags. For example, `*italic` should be rendered as `<p>*italic</p>`. - The order of precedence for inline elements is `***`, then `**`, then `*`. - Assume input is a single multi-line string. - Do not implement support for any other Markdown features like links, images, blockquotes, or ordered lists. - The output HTML does not need to be a full document (no `<html>` or `<body>` tags are required). **Example Input:** ```markdown # Header 1 This is a paragraph with **bold** and *italic* text. This is the same paragraph. - List item one - List item two with ***bold and italic*** - Nested list item - Back to the first level ```python def hello(): print("Hello, World!") ``` ```

402
Apr 22, 2026 09:40

Coding

Anthropic Claude Sonnet 4.6 VS OpenAI GPT-5.4

Implement a Thread-Safe Token Bucket Rate Limiter in Python

Write a Python class named `TokenBucketRateLimiter` that implements the token bucket algorithm for rate limiting. The implementation must be thread-safe and should not use any external libraries for state management (like Redis). The class should have the following specifications: 1. An `__init__(self, capacity, refill_rate)` method: * `capacity`: The maximum number of tokens the bucket can hold. * `refill_rate`: The number of tokens that are added to the bucket per second. 2. A `consume(self, tokens)` method: * This method attempts to consume a given number of `tokens` from the bucket. * It should return `True` if the tokens can be consumed successfully, and `False` otherwise. * The bucket should be refilled with tokens based on the time elapsed since the last call before attempting to consume. 3. Thread Safety: * The class must be safe to use from multiple concurrent threads. All operations that modify the bucket's state (like refilling and consuming tokens) must be atomic. Provide the complete class implementation with necessary imports.

379
Apr 16, 2026 09:37

Related Links

X f L