Orivel Orivel
Open menu

Summarize the Lantern Loop Night-Transit Pilot

Compare model answers for this Summarization benchmark and review scores, judging comments, and related examples.

Login or register to use likes and favorites. Register

X f L

Contents

Task Overview

Benchmark Genres

Summarization

Task Creator Model

Answering Models

Judge Models

Task Prompt

Read the fictional municipal briefing below and write a 180–230 word executive summary in prose. The summary must preserve: the pilot’s purpose and design; the most important ridership, cost, safety, and employment findings; the principal limitations on interpreting the evidence; the competing stakeholder views; and the review team’s recommendation, including its conditions and funding implication. Clearly distinguish measured results from estimates or self-reported outcomes. Do not introduce facts, calculations, o...

Show more

Read the fictional municipal briefing below and write a 180–230 word executive summary in prose. The summary must preserve: the pilot’s purpose and design; the most important ridership, cost, safety, and employment findings; the principal limitations on interpreting the evidence; the competing stakeholder views; and the review team’s recommendation, including its conditions and funding implication. Clearly distinguish measured results from estimates or self-reported outcomes. Do not introduce facts, calculations, or recommendations absent from the passage.

Source passage:

In March 2025, the City of Bellwether began a six-month night-transit experiment called the Lantern Loop. The project responded to complaints from hospital staff, hospitality workers, warehouse employees, and students who said that regular bus service ended before many late shifts did. Before the experiment, Bellwether’s last scheduled buses left the central interchange at 11:20 p.m.; afterward, most people without cars relied on taxis, informal rides, or walks of up to four kilometers. The pilot was intended to test whether a limited overnight network could provide useful access without committing the city to a permanent, citywide service. It was not designed to replace daytime routes or operate at the same frequency.

The Lantern Loop consisted of two circular routes running in opposite directions between midnight and 4:30 a.m., Thursday through Sunday. Each loop connected the central interchange with Northbank Hospital, the Arlen warehouse district, East Quay’s restaurant corridor, and two neighborhoods with high numbers of shift workers. Buses arrived at major stops approximately every 45 minutes. The standard fare was 2 crowns, compared with 3 crowns during the day, and riders transferring from the final evening buses paid nothing extra. The city used four older diesel buses already in its reserve fleet rather than buying new vehicles. Stops were fitted with brighter lighting and temporary emergency-call buttons, while two transit stewards circulated between buses instead of assigning one steward to every vehicle.

The council approved a maximum pilot budget of 780,000 crowns. Final direct spending was 692,000 crowns: 318,000 for drivers and stewards, 166,000 for fuel and maintenance, 121,000 for stop lighting and call buttons, and 87,000 for administration, promotion, and evaluation. Fare revenue totaled 94,000 crowns, leaving a net municipal cost of 598,000 crowns. The finance office noted that the lighting equipment could remain in use for several years, although the temporary call-button system would require a new contract if the service continued. The pilot’s average net subsidy was 7.18 crowns per recorded passenger trip. For comparison, the city reports a systemwide subsidy of 4.90 crowns per trip, but that figure combines crowded peak services with quieter routes and is therefore not a direct measure of whether the night service was inefficient.

Automated counters recorded 83,240 passenger trips during the six months. Monthly use rose from 10,180 trips in March to 16,070 in August, though part of the increase coincided with warmer weather and the summer festival season. Thursday nights were consistently the quietest, averaging 61 passengers per service hour across both loops, while Saturday nights averaged 104. The busiest stop was Northbank Hospital, which accounted for 27 percent of boardings. Arlen district stops accounted for 21 percent, East Quay for 18 percent, the two residential areas together for 29 percent, and all other stops for 5 percent. Crowding occurred on 14 Saturday departures, but most buses had spare seats. On-time performance was 88 percent, below the daytime network’s 92 percent, mainly because street-cleaning closures forced overnight detours.

Safety results were mixed but generally favorable. Transit security logs recorded nine incidents on buses or at pilot stops: six verbal disputes, two cases of property damage, and one minor assault that did not require hospital treatment. No driver was physically attacked. During the comparable Thursday-to-Sunday overnight periods in the same areas a year earlier, police had recorded 15 incidents near the relevant stops, including three assaults. However, the review team warned that the two sets of records were compiled differently and that police reports cannot establish how many incidents involved people who would have used the bus. A rider survey found that 74 percent of respondents felt safer traveling at night because of the service. That result reflects perceptions among survey participants, not a measured reduction in crime.

To examine employment effects, evaluators surveyed 1,200 riders by text message, receiving 486 complete responses. Of those respondents, 112 said the Lantern Loop had allowed them to accept extra shifts, and 38 said it had helped them take a new job. Employers at Northbank Hospital and three East Quay restaurants separately reported fewer late-shift absences, but only the hospital supplied payroll records. Those records showed that unplanned absences on eligible night shifts fell by 11 percent compared with the same months in 2024. Hospital managers also introduced a stricter attendance policy in May, so evaluators could not determine how much of the improvement resulted from transit access. The warehouse association declined to share company-level attendance data, citing confidentiality concerns.

The pilot did not benefit all areas equally. Residents of western Bellwether argued that the route map favored major institutions and eastern neighborhoods. A community group proposed extending one loop six kilometers west to serve the Brindle Estate, where car ownership is low. Transit planners estimated that the extension would add 14 minutes to each circuit, making the advertised 45-minute interval unreliable unless a fifth bus and another driver were added. Disability advocates praised the low-floor buses but documented 23 occasions when temporary construction barriers made boarding areas difficult to reach. Three of those barriers remained unresolved for more than a week. The public works department has since assigned a named inspector to overnight-stop accessibility complaints.

Environmental claims also require qualification. Because the reserve buses were diesel vehicles, the pilot produced an estimated 126 metric tons of carbon-dioxide-equivalent emissions. The sustainability office modeled that riders would otherwise have generated about 91 metric tons through taxi, private-car, and ride-hailing trips, based on survey answers about previous travel habits. The resulting estimated net increase was therefore 35 metric tons. Yet the model did not account for people who previously declined trips altogether, and self-reported travel habits may be inaccurate. Replacing the reserve fleet with four leased electric buses would reduce operating emissions, but preliminary supplier quotes indicate an additional annual lease cost of 240,000 crowns, excluding charging equipment.

Stakeholders interpreted the evidence differently. The Chamber of Evening Commerce called the pilot an economic-access program rather than a transport expense and requested nightly service, including Mondays through Wednesdays. The drivers’ union supported continuation only if overnight shifts remained voluntary and included the current 18 percent wage premium. A taxpayers’ association argued that the subsidy per trip was too high and recommended subsidized taxi vouchers for verified workers instead. Evaluators cautioned that no taxi-voucher trial had been conducted, so its cost, availability, and effect on riders could not yet be compared reliably with the bus service. Rider groups favored retaining the low fare and opposed restricting access to people who could prove employment.

The review team recommends extending the Lantern Loop for twelve months, but not yet making it permanent or expanding it to seven nights a week. Under the recommendation, the existing Thursday-to-Sunday schedule and 2-crown fare would remain, while Friday and Saturday frequency would improve from 45 to 30 minutes between 12:30 and 2:30 a.m. The city would lease one additional conventional bus for those peak periods, add a steward, correct all documented access barriers, and run a small taxi-voucher comparison in the western districts. Continued service should be conditional on quarterly reporting of ridership, cost per trip, accessibility failures, incidents, and on-time performance. The team estimates a twelve-month net municipal cost of 1.34 million crowns. Only 900,000 crowns is available in the existing transit allocation, so approval would require either 440,000 crowns in new funding or reductions elsewhere. A decision is scheduled for the council’s 18 October budget meeting.

Judging Policy

A strong answer will synthesize rather than list every detail, remain within 180–230 words, and accurately convey the service design, central quantitative findings, evidentiary caveats, stakeholder disagreement, and conditional twelve-month recommendation. It should preserve the distinction between recorded data, modeled estimates, institutional reports, and rider self-reports. The funding gap and proposed operational changes should be stated correctly. Minor details may be omitted, but the answer should not distor...

Show more

A strong answer will synthesize rather than list every detail, remain within 180–230 words, and accurately convey the service design, central quantitative findings, evidentiary caveats, stakeholder disagreement, and conditional twelve-month recommendation. It should preserve the distinction between recorded data, modeled estimates, institutional reports, and rider self-reports. The funding gap and proposed operational changes should be stated correctly. Minor details may be omitted, but the answer should not distort comparisons, imply causation where the passage does not establish it, or add unsupported conclusions.

Answers

Between March and August 2025, Bellwether piloted the Lantern Loop, a six-month overnight bus service designed to address gaps left by daytime routes ending at 11:20 p.m., without committing to a permanent citywide network. Two circular routes ran Thursday–Sunday, midnight to 4:30 a.m., using four existing diesel buses, connecting the hospital, warehouse district, restaurant corridor, and shift-worker neighborhoods, with fares at 2 crowns.

Measured results: 83,240 recorded trips, rising from 10,180 (March) to 16,070 (August), partly due to seasonal effects. Net municipal cost was 598,000 crowns (7.18 crowns/trip...

Show Full Answer

Between March and August 2025, Bellwether piloted the Lantern Loop, a six-month overnight bus service designed to address gaps left by daytime routes ending at 11:20 p.m., without committing to a permanent citywide network. Two circular routes ran Thursday–Sunday, midnight to 4:30 a.m., using four existing diesel buses, connecting the hospital, warehouse district, restaurant corridor, and shift-worker neighborhoods, with fares at 2 crowns.

Measured results: 83,240 recorded trips, rising from 10,180 (March) to 16,070 (August), partly due to seasonal effects. Net municipal cost was 598,000 crowns (7.18 crowns/trip subsidy), against 692,000 crowns spent and 94,000 in fare revenue. On-time performance was 88%, below the daytime network's 92%. Security logs recorded nine incidents, fewer than the prior year's 15 police-recorded incidents, though the two records aren't directly comparable. A survey found 74% of riders felt safer—a perception, not a measured crime reduction. Employment effects were self-reported: 112 of 486 survey respondents said the service enabled extra shifts; hospital absenteeism fell 11%, but a concurrent policy change confounds this estimate. Estimated net emissions increase was 35 tons CO2e, based on modeled, self-reported prior travel habits.

Stakeholders diverged: commerce groups sought nightly expansion; the drivers' union wanted voluntary shifts with wage premiums; taxpayers favored taxi vouchers (untested); riders opposed employment-based access restrictions.

The review team recommends a 12-month extension with increased weekend frequency, an added bus, accessibility fixes, and a voucher trial, conditioned on quarterly reporting. Estimated cost is 1.34 million crowns, exceeding the 900,000-crown allocation by 440,000 crowns, pending an 18 October council decision.

Result

#1 | Winner

Winning Votes

2 / 3

Average Score

78

Total Score

74

Overall Comments

Answer A is an exceptionally faithful and comprehensive summary: it reproduces every key figure correctly, consistently distinguishes recorded data from perceptions, estimates, and confounded results, and captures all conditions of the recommendation including quarterly reporting, the funding gap against the existing allocation, and the decision date. Its main failures are exceeding the mandatory 180–230 word limit by roughly 23 words and a slightly telegraphic, parenthetical-heavy style, plus a minor mislabeling of measured payroll data under a 'self-reported' heading.

View Score Details

Faithfulness

Weight 40%
83

Answer A is highly accurate: all figures (83,240 trips, 598,000 net cost, 7.18 subsidy, 88% vs 92% on-time, 9 vs 15 incidents, 35 tons CO2e, 1.34M cost, 440,000 gap) match the source, and it explicitly flags perception vs measurement, confounded absenteeism data, and modeled emissions. The only flaw is placing the hospital's 11% payroll-measured drop under the heading 'employment effects were self-reported,' a minor mislabel partially corrected by the confounding caveat.

Coverage

Weight 20%
85

Answer A covers every required element: purpose and design, ridership with seasonal caveat, full cost breakdown, safety with comparability caveat, employment findings including the 112/486 self-reports and the confounded payroll data, on-time performance, emissions estimate, all four stakeholder positions, and the full recommendation with quarterly-reporting condition, funding gap against the 900,000 allocation, and the 18 October decision date.

Compression

Weight 15%
40

Answer A runs to roughly 253 words, clearly exceeding the mandatory 180–230 word limit by about 10%. While the density of information per word is high and there is little redundancy, violating an explicit hard constraint of the task is a significant compression failure.

Clarity

Weight 15%
70

Answer A is readable but dense and somewhat telegraphic: the 'Measured results:' colon construction and heavy parentheticals ('(7.18 crowns/trip subsidy)', '(untested)') make it feel like compressed notes rather than fluid executive prose, though every sentence is unambiguous.

Structure

Weight 10%
74

Answer A follows a logical arc from design to measured results to stakeholder views to recommendation, mirroring the briefing's priorities. Limitations are woven into the findings rather than gathered separately, which works but makes the caveat layer slightly harder to isolate.

Total Score

81

Overall Comments

Answer A is a highly faithful and comprehensive summary that accurately captures all the key details and nuances from the source text. It correctly distinguishes between different types of evidence. However, its major weakness is its failure to adhere to the word count limit; at 254 words, it significantly exceeds the 230-word maximum, which is a critical failure in a compression task. Its structure is logical, but less elegant than Answer B's.

View Score Details

Faithfulness

Weight 40%
95

The summary is extremely faithful, capturing all key figures and nuances from the source text without introducing errors or distortions.

Coverage

Weight 20%
90

The answer covers all the required points from the prompt, including the pilot's design, findings, limitations, stakeholder views, and recommendation.

Compression

Weight 15%
40

The answer fails on this criterion. At 254 words, it is significantly over the 230-word maximum, which is a critical flaw for a summarization task.

Clarity

Weight 15%
75

The summary is generally clear, but the use of an explicit label ('Measured results:') within the prose is slightly unconventional and breaks the flow.

Structure

Weight 10%
75

The structure is logical, with paragraphs dedicated to different aspects of the summary. The overall organization is effective.

Judge Models OpenAI GPT-5.6

Total Score

78

Overall Comments

Answer A is substantively strong and carefully distinguishes recorded, surveyed, modeled, and noncomparable evidence. It covers the central ridership, cost, safety, employment, emissions, stakeholder, and funding findings, including quarterly reporting. Its main weakness is substantial noncompliance with the 180–230-word limit, along with omission of some recommendation specifics and the reported new-job outcome.

View Score Details

Faithfulness

Weight 40%
87

Accurately reports the principal figures and consistently labels measured, surveyed, modeled, and noncomparable evidence. Minor weaknesses include calling the confounded payroll result an employment estimate and compressing some operational details.

Coverage

Weight 20%
82

Covers purpose, core design, ridership, cost, reliability, safety, self-reported employment, emissions, stakeholder disagreement, recommendation, reporting conditions, and funding gap. It omits the 38 reported new jobs and several precise recommendation details, including the retained fare and exact 30-minute period.

Compression

Weight 15%
45

The summary is information-dense and avoids a list-like presentation, but it is approximately 253 words, materially exceeding the required 180–230-word range.

Clarity

Weight 15%
80

Clear prose and effective signals such as “Measured results,” “self-reported,” and “estimated” make evidentiary status easy to follow. The dense second paragraph slightly reduces readability.

Structure

Weight 10%
78

Uses a logical sequence from purpose and design through measured findings, stakeholder views, and recommendation. The long findings paragraph is somewhat overloaded.

The Lantern Loop night-transit pilot, running March-August 2025, aimed to provide limited overnight access for late-shift workers and students, testing viability without committing to a permanent citywide service. Operating Thursday-Sunday from midnight to 4:30 a.m. on two circular routes, it used existing diesel buses and offered a 2-crown fare.

Measured ridership totaled 83,240 trips, increasing monthly, though partly coinciding with warmer weather. The net municipal cost was 598,000 crowns, with a 7.18 crowns per trip subsidy, higher than the systemwide average but not directly comparable. Safety logs recorde...

Show Full Answer

The Lantern Loop night-transit pilot, running March-August 2025, aimed to provide limited overnight access for late-shift workers and students, testing viability without committing to a permanent citywide service. Operating Thursday-Sunday from midnight to 4:30 a.m. on two circular routes, it used existing diesel buses and offered a 2-crown fare.

Measured ridership totaled 83,240 trips, increasing monthly, though partly coinciding with warmer weather. The net municipal cost was 598,000 crowns, with a 7.18 crowns per trip subsidy, higher than the systemwide average but not directly comparable. Safety logs recorded nine incidents, fewer than police reports in comparable prior periods, but data compilation differed. A rider survey reported 74% felt safer. Hospital payroll records showed an 11% drop in night-shift absences, but a new attendance policy also contributed. Environmental analysis estimated a net increase of 35 metric tons of CO2e emissions, based on self-reported travel habits.

Limitations include external factors influencing ridership, differing safety data, and the hospital's policy change. Stakeholders held varied views: businesses sought expanded service, the union wanted voluntary shifts with premiums, taxpayers favored taxi vouchers, and riders opposed restricted access.

The review team recommends a 12-month extension, not permanent, maintaining the current schedule and fare, but improving Friday/Saturday frequency. Conditions include leasing an additional bus, adding a steward, correcting access barriers, and trialing taxi vouchers. The estimated 1.34 million crown cost requires 440,000 crowns in new funding.

Result

#2

Winning Votes

1 / 3

Average Score

80

Total Score

74

Overall Comments

Answer B is a well-organized, fluent summary that stays within the word limit and includes a dedicated limitations sentence. However, it introduces a mild causal overstatement (the attendance policy 'also contributed' where the passage says attribution could not be determined), overreads the ridership trend as 'increasing monthly,' and omits several required elements: the quarterly-reporting condition, the rider employment self-reports, on-time performance, and the funding context of the 900,000-crown allocation and the October decision date.

View Score Details

Faithfulness

Weight 40%
72

Answer B is mostly accurate but contains two mild distortions: saying the new attendance policy 'also contributed' asserts a causal contribution the passage explicitly says evaluators could not determine, and 'increasing monthly' overreads the March-to-August rise as a month-over-month pattern. It also softens the perception caveat on the 74% safety figure compared with the source's explicit warning, though 'survey reported' partially preserves the distinction.

Coverage

Weight 20%
68

Answer B covers purpose, design, ridership, cost, safety, absenteeism, emissions, stakeholders, and the core recommendation, and adds a useful explicit limitations sentence. However, it omits the quarterly-reporting condition (an explicit condition the prompt requires), the employment survey self-reports (112 riders/extra shifts), on-time performance, the 900,000-crown allocation context, the 'or reductions elsewhere' alternative, and the decision date.

Compression

Weight 15%
82

Answer B sits at roughly 229 words, just inside the required range, and achieves genuine synthesis: it groups limitations into a single sentence, condenses stakeholder views into one clause per group, and avoids enumerating minor details, which is exactly the compression behavior the judging policy rewards.

Clarity

Weight 15%
76

Answer B reads as smoother, more natural executive prose with well-connected sentences and clear signposting ('Limitations include...', 'Conditions include...'). A few caveats are stated a bit flatly, but the overall flow is easier for a decision-maker to absorb quickly.

Structure

Weight 10%
75

Answer B uses a clean four-part structure (design, findings, limitations plus stakeholders, recommendation) with a dedicated limitations sentence that directly answers the prompt's requirement to identify interpretive constraints, giving it a marginally clearer organizational scaffold.

Total Score

90

Overall Comments

Answer B is an excellent summary that successfully balances detail with conciseness. It covers all the required elements of the prompt faithfully and accurately. Crucially, it adheres almost perfectly to the word count constraint, coming in at 231 words. Its structure is particularly effective, with a dedicated paragraph for the pilot's limitations, which enhances clarity. The prose is clear and professional.

View Score Details

Faithfulness

Weight 40%
95

The summary is extremely faithful, accurately representing all the core findings, caveats, and recommendations from the source document.

Coverage

Weight 20%
90

The answer provides excellent coverage of all prompt requirements, addressing the pilot's purpose, results, limitations, stakeholder positions, and the final recommendation.

Compression

Weight 15%
85

The answer demonstrates excellent compression. At 231 words, it is only one word over the maximum limit, successfully condensing a large amount of information into the required length.

Clarity

Weight 15%
85

The summary is very clear and easy to understand. The decision to dedicate a separate paragraph to the study's limitations enhances the overall clarity of the argument.

Structure

Weight 10%
85

The structure is excellent and highly logical. The flow from findings to limitations, then to stakeholder views and the final recommendation, is very effective and easy to follow.

Judge Models OpenAI GPT-5.6

Total Score

76

Overall Comments

Answer B is concise, readable, well organized, and within the requested length. It captures most headline findings and major stakeholder positions, but omits the rider-reported employment outcomes and the recommendation’s quarterly reporting condition. It also overstates the attendance policy’s contribution and incorrectly presents new funding as the only way to close the funding gap.

View Score Details

Faithfulness

Weight 40%
75

Most figures are accurate and evidentiary caveats are generally retained. However, saying the attendance policy “also contributed” asserts an effect the evaluators could not determine, and saying the cost “requires” new funding omits the stated alternative of reductions elsewhere.

Coverage

Weight 20%
66

Includes the main design and headline ridership, cost, safety, payroll, environmental, stakeholder, and recommendation points. Important omissions include both rider-reported employment outcomes, quarterly performance reporting, and the alternative of budget reductions; design and operational changes are also less specific.

Compression

Weight 15%
85

Fits the required word range and efficiently synthesizes a long briefing without becoming choppy. Its concision does, however, come at the cost of several required findings and conditions.

Clarity

Weight 15%
80

Direct, polished prose with clear transitions and accessible presentation of quantitative findings. Some phrasing blurs uncertainty, especially the claim that the attendance policy contributed to the decline.

Structure

Weight 10%
77

Follows a coherent executive-summary progression from design to findings, limitations, stakeholders, and recommendation. The separate limitations sentence is useful but partly repetitive, and recommendation conditions are incompletely framed.

Comparison Summary

Final rank order is determined by judge-wise rank aggregation (average rank + Borda tie-break). Average score is shown for reference.

Judges: 3

Winning Votes

2 / 3

Average Score

78
View this answer

Winning Votes

1 / 3

Average Score

80
View this answer

Judging Results

Judge Models OpenAI GPT-5.6

Why This Side Won

Answer A wins because the heavily weighted faithfulness and coverage criteria favor its more precise treatment of evidence and fuller account of the conditional recommendation. In particular, it preserves the self-reported employment finding, explicitly distinguishes perceived safety from measured crime reduction, notes the untested status of taxi vouchers, and includes quarterly reporting. These advantages outweigh its significant word-limit violation and weaker compression.

Why This Side Won

Answer B is the winner because it successfully meets all the task's requirements, most notably the strict word count. While both answers are highly faithful and comprehensive, Answer A significantly exceeds the 230-word limit, which is a major failure for a summarization task. Answer B demonstrates superior compression skills by conveying the necessary information within the specified length. Additionally, B has a slightly clearer and more logical structure.

Why This Side Won

Answer A wins on the weighted result. Faithfulness carries 40% of the weight, and A is markedly stronger there: it preserves every measured-versus-estimated distinction the prompt demands and contains no causal overstatements, while B asserts the attendance policy 'contributed' despite the passage saying attribution could not be determined, and overreads the ridership trend. A also clearly wins Coverage (20%), including the quarterly-reporting condition, employment self-report figures, on-time performance, and full funding context that B omits. B wins Compression decisively because A exceeds the 230-word cap, and B is marginally better on Clarity and Structure, but those three criteria together (40%) are not enough to offset A's advantages on the two most heavily weighted criteria. The weighted totals implied by the criterion scores place A narrowly but consistently ahead.

X f L