| BIZ-1 | AI-draft reply release gate | infra_ops | 20 | 0 |
| BIZ-10 | Language and native-speaker routing | speech_dialogue | 22 | 0 |
| BIZ-100 | Supplier risk event tagging | finance_insurance | 433 | 0 |
| BIZ-101 | SOX control evidence attestation check | legal_compliance | 42 | 0 |
| BIZ-102 | FNOL triage: line, severity band, fast-track | finance_insurance | 22,836 | 10 |
| BIZ-103 | Fraud indicator scoring for SIU referral | security_safety | 18,042 | 3 |
| BIZ-104 | Subrogation opportunity detection | finance_insurance | 13,146 | 4 |
| BIZ-105 | Medical bill review for injury claims (relatedness and duplicates) | medicine_health | 2 | 0 |
| BIZ-106 | Commercial submission triage (appetite and completeness) | finance_insurance | 1 | 0 |
| BIZ-107 | Policy-change request classification to closed actions | finance_insurance | 876 | 1 |
| BIZ-108 | Claim document indexing and missing-document checklist | finance_insurance | 13 | 0 |
| BIZ-109 | Litigation and attorney-representation propensity | legal_compliance | 8 | 6 |
| BIZ-11 | Auto-close and pending-state decisions | support_ops | 12,994 | 18 |
| BIZ-110 | Coverage pre-determination against exclusion list | finance_insurance | 3,496 | 5 |
| BIZ-111 | Regulatory-complaint risk detection in correspondence | legal_compliance | 20,473 | 5 |
| BIZ-112 | Repair estimate and supplement review | math_reasoning | 13,144 | 4 |
| BIZ-113 | Property damage photo severity (successor) | travel_hospitality | 16 | 2 |
| BIZ-114 | Life/disability claim intake classification | finance_insurance | 2 | 0 |
| BIZ-115 | Travel and health claim line eligibility | medicine_health | 9,431 | 6 |
| BIZ-116 | Broker email intent to workflow | language_nlp | 404 | 2 |
| BIZ-117 | Reserve-adequacy signal from adjuster notes | finance_insurance | 408 | 0 |
| BIZ-118 | AML alert level-1 disposition | finance_insurance | 244 | 1 |
| BIZ-119 | Sanctions / PEP name-screening hit adjudication | finance_insurance | 10,832 | 6 |
| BIZ-12 | 100%-coverage QA scoring of agent transcripts | finance_insurance | 1,182 | 3 |
| BIZ-120 | KYC document classification and completeness | finance_insurance | 1,660 | 10 |
| BIZ-121 | Adverse-media relevance and severity | finance_insurance | 7,168 | 5 |
| BIZ-122 | Payment purpose coding from reference text | finance_insurance | 0 | 0 |
| BIZ-123 | Loan-file document indexing and consistency flags | finance_insurance | 9,198 | 2 |
| BIZ-124 | Small-business loan pre-screen from narrative | finance_insurance | 673 | 1 |
| BIZ-125 | Complaint classification for regulatory reporting | legal_compliance | 49,871 | 18 |
| BIZ-126 | Card/bank dispute intake and provisional-credit decision | finance_insurance | 9,350 | 0 |
| BIZ-127 | Collections hardship and vulnerability detection with next action | security_safety | 20,471 | 5 |
| BIZ-128 | Scam-victim and social-engineering detection in contact-center transcripts | security_safety | 3,231 | 0 |
| BIZ-129 | Mortgage condition clearing | finance_insurance | 34 | 0 |
| BIZ-13 | Reopen-risk prediction at close | support_ops | 12,541 | 13 |
| BIZ-130 | Covenant and credit-memo signal reading | finance_insurance | 288 | 1 |
| BIZ-131 | Business account opening: industry and high-risk detection | business_ops | 6,789 | 10 |
| BIZ-132 | Trade-finance document discrepancy check (successor) | finance_insurance | 246 | 3 |
| BIZ-133 | Adverse-action reason selection | finance_insurance | 9,199 | 2 |
| BIZ-134 | Clause-level playbook deviation classification | legal_compliance | 34,901 | 12 |
| BIZ-135 | NDA triage: sign-as-is / negotiate / escalate | legal_compliance | 19,779 | 8 |
| BIZ-136 | Third-party paper risk tiering and reviewer assignment | business_ops | 1,335 | 5 |
| BIZ-137 | Obligation and renewal-term presence flags for abstraction | legal_compliance | 19,596 | 8 |
| BIZ-138 | eDiscovery relevance and privilege first pass | legal_compliance | 6 | 6 |
| BIZ-139 | Regulatory-change applicability screening | legal_compliance | 13,346 | 14 |
| BIZ-14 | App-store / marketplace review triage | commerce_marketing | 10,392 | 8 |
| BIZ-140 | DSAR intake and personal-data classification | legal_compliance | 10 | 1 |
| BIZ-141 | Regulated-claims review in marketing (financial promotions, health claims) | medicine_health | 13 | 0 |
| BIZ-142 | Ethics-hotline report triage | support_ops | 14 | 1 |
| BIZ-143 | Litigation-hold custodian relevance | legal_compliance | 7 | 1 |
| BIZ-144 | Trademark watch: confusing-similarity screening | legal_compliance | 16 | 1 |
| BIZ-145 | In-house legal intake routing | legal_compliance | 877 | 6 |
| BIZ-146 | Gifts, entertainment and conflict disclosure review | finance_insurance | 3,581 | 2 |
| BIZ-147 | Court notice and filing classification for docketing | legal_compliance | 23 | 0 |
| BIZ-148 | Redline acceptability judgment | business_ops | 19,556 | 8 |
| BIZ-149 | Vendor security questionnaire answer sufficiency | finance_insurance | 8 | 0 |
| BIZ-15 | Effort / complexity estimation for capacity planning | math_reasoning | 8 | 4 |
| BIZ-150 | Prior-authorization criteria checklist | medicine_health | 703 | 2 |
| BIZ-151 | Referral and fax intake classification | medicine_health | 1,198 | 1 |
| BIZ-152 | Coding support: candidate-code confirmation and E/M level | commerce_marketing | 408 | 1 |
| BIZ-153 | Denial triage and appeal worthiness | business_ops | 3,322 | 5 |
| BIZ-154 | Patient-portal message triage | medicine_health | 13 | 0 |
| BIZ-155 | Nurse-line protocol routing (non-diagnostic disposition) | medicine_health | 238 | 2 |
| BIZ-156 | CDI query candidate detection | medicine_health | 1,193 | 1 |
| BIZ-157 | Appointment type selection and no-show risk | medicine_health | 1,180 | 1 |
| BIZ-158 | Eligibility-response discrepancy classification | business_ops | 16 | 0 |
| BIZ-159 | Patient-safety event report classification | medicine_health | 20 | 0 |
| BIZ-16 | Engineering escalation gate | science_eng | 4,557 | 10 |
| BIZ-160 | Clinical-trial pre-screening checklist | medicine_health | 9 | 0 |
| BIZ-161 | Misdirected-PHI detection in outbound communications | security_safety | 0 | 6 |
| BIZ-162 | Charge-capture missed-charge detection | medicine_health | 13 | 0 |
| BIZ-163 | Provider credentialing document verification completeness | security_safety | 10 | 1 |
| BIZ-164 | Level-of-care criteria checklist for utilization review | medicine_health | 627 | 2 |
| BIZ-165 | Patient complaint and grievance classification | medicine_health | 31 | 0 |
| BIZ-166 | Listing fair-housing and disclosure compliance | legal_compliance | 7 | 0 |
| BIZ-167 | Real-estate lead intent and qualification | travel_hospitality | 3 | 0 |
| BIZ-168 | Rental application document consistency and fraud signals | security_safety | 9 | 0 |
| BIZ-169 | Maintenance request triage | business_ops | 394 | 4 |
| BIZ-17 | Post-call disposition and compliance coding | legal_compliance | 29 | 0 |
| BIZ-170 | Lease abstraction into closed fields | travel_hospitality | 631 | 1 |
| BIZ-171 | Inspection report defect coding (successor) | business_ops | 3,367 | 7 |
| BIZ-172 | Title and closing document classification and checklist | travel_hospitality | 20 | 0 |
| BIZ-173 | Tenant communication routing and legal-risk flags | legal_compliance | 15,059 | 12 |
| BIZ-174 | Comparable-property similarity judgment | travel_hospitality | 5,477 | 6 |
| BIZ-175 | HOA violation report classification | travel_hospitality | 5,053 | 11 |
| BIZ-176 | Mortgage-servicing correspondence classification | finance_insurance | 1,931 | 0 |
| BIZ-177 | Listing photo tagging (successor) | travel_hospitality | 16 | 2 |
| BIZ-178 | Shipment exception classification and next action | code_software | 3,684 | 3 |
| BIZ-179 | Freight invoice accessorial validity | finance_insurance | 11 | 0 |
| BIZ-18 | KB-article answerability judge (deflection) | support_ops | 29 | 0 |
| BIZ-180 | Cargo claim intake and liability triage | finance_insurance | 3,704 | 5 |
| BIZ-181 | Customs classification (HS code) and restricted-goods flags | legal_compliance | 1,482 | 4 |
| BIZ-182 | Load tender and carrier email to closed actions | business_ops | 6 | 0 |
| BIZ-183 | Proof-of-delivery exception check (successor) | code_software | 22 | 0 |
| BIZ-184 | Supplier order-acknowledgment discrepancy detection | finance_insurance | 4 | 2 |
| BIZ-185 | Receiving discrepancy coding | business_ops | 321 | 0 |
| BIZ-186 | Fleet DVIR defect severity and incident report classification | science_eng | 13,146 | 4 |
| BIZ-187 | Dispatch exception action selection | code_software | 13 | 0 |
| BIZ-188 | Address and location deliverability judgment | geo_place | 146 | 1 |
| BIZ-189 | Returns (RMA) disposition from inspection notes | commerce_marketing | 788 | 0 |
| BIZ-19 | Systemic-issue early warning | business_ops | 623 | 0 |
| BIZ-190 | Telematics event triage (successor) | finance_insurance | 20 | 0 |
| BIZ-191 | Supplier communication risk signals | finance_insurance | 4 | 0 |
| BIZ-192 | Multi-property guest review coding and response prioritization | travel_hospitality | 22 | 1 |
| BIZ-193 | Guest message routing and upsell detection | travel_hospitality | 4 | 6 |
| BIZ-194 | Booking change/cancellation request to fare-rule action | finance_insurance | 10,361 | 4 |
| BIZ-195 | Traveler risk / duty-of-care triage | travel_hospitality | 2,315 | 11 |
| BIZ-196 | Service-recovery compensation tier | finance_insurance | 12 | 0 |
| BIZ-197 | Group / event RFP triage for hotels | travel_hospitality | 11 | 0 |
| BIZ-198 | Short-term-rental guest screening and party risk | travel_hospitality | 559 | 3 |
| BIZ-199 | Food-safety and incident report classification | infra_ops | 9,199 | 6 |
| BIZ-2 | Duplicate / merge ticket judgment | support_ops | 22 | 0 |
| BIZ-20 | Inbound lead qualification and disqualification | commerce_marketing | 6,438 | 10 |
| BIZ-200 | Dietary and special-request handling | business_ops | 338 | 3 |
| BIZ-201 | Delay-cause classification for compensation eligibility | finance_insurance | 13,605 | 11 |
| BIZ-202 | Rubric-based short-answer and essay scoring | education | 13,319 | 10 |
| BIZ-203 | Student support routing with safety flags | education | 394 | 3 |
| BIZ-204 | Admissions application triage | education | 9 | 1 |
| BIZ-205 | Academic-integrity signal triage | education | 393 | 11 |
| BIZ-206 | Content-to-standards alignment tagging | business_ops | 6 | 0 |
| BIZ-207 | Forum moderation and at-risk detection | security_safety | 111,849 | 72 |
| BIZ-208 | Early-alert from advisor notes and messages | infra_ops | 13 | 0 |
| BIZ-209 | Accommodation documentation classification | legal_compliance | 6 | 6 |
| BIZ-21 | Lead-to-rep routing by territory, segment and product | business_ops | 388 | 1 |
| BIZ-210 | Transfer-credit equivalency judgment | finance_insurance | 406 | 6 |
| BIZ-211 | K-12 family communication classification | education | 6 | 0 |
| BIZ-212 | Public-records / FOIA request triage | legal_compliance | 2,875 | 1 |
| BIZ-213 | 311 service-request classification and duplicate detection | gov_civic | 15,059 | 13 |
| BIZ-214 | Permit application completeness and review track | gov_civic | 2,070 | 10 |
| BIZ-215 | Benefits application document classification and eligibility pre-check | finance_insurance | 15 | 0 |
| BIZ-216 | Rulemaking public-comment coding | gov_civic | 659 | 2 |
| BIZ-217 | Constituent correspondence classification | gov_civic | 2,635 | 6 |
| BIZ-218 | Grant application eligibility screening | finance_insurance | 2,500 | 8 |
| BIZ-219 | Code-enforcement complaint triage | business_ops | 813 | 10 |
| BIZ-22 | Pipeline-stage suggestion from activity text | infra_ops | 7 | 0 |
| BIZ-220 | E-filing clerk review | legal_compliance | 3,515 | 0 |
| BIZ-221 | Tax-agency correspondence classification | finance_insurance | 2 | 0 |
| BIZ-222 | Donor correspondence and gift-intent classification | finance_insurance | 12 | 6 |
| BIZ-223 | Grant-opportunity fit screening | finance_insurance | 1,020 | 4 |
| BIZ-224 | Volunteer application screening and role matching | hr_workforce | 15 | 0 |
| BIZ-225 | Program intake and case-management triage | code_software | 1,478 | 1 |
| BIZ-226 | Crisis-line message risk triage (human-supervised) | security_safety | 1,416 | 2 |
| BIZ-227 | Restricted-fund allowability of expenses | finance_insurance | 538 | 0 |
| BIZ-228 | Safeguarding and consent review of communications assets | legal_compliance | 6 | 0 |
| BIZ-229 | Legislation relevance tagging for advocacy | llm_rag_data | 25,267 | 13 |
| BIZ-23 | Deal-risk and slippage scoring | finance_insurance | 0 | 1 |
| BIZ-24 | Outbound reply classification for sequences | commerce_marketing | 26 | 1 |
| BIZ-25 | Objection and competitor tagging on call transcripts | infra_ops | 3 | 0 |
| BIZ-26 | Account / contact deduplication judgment | data_quality | 10,687 | 8 |
| BIZ-27 | Buying-role / persona classification | hr_workforce | 2,206 | 4 |
| BIZ-28 | Expansion and churn signals from support/usage notes | finance_insurance | 405 | 2 |
| BIZ-29 | RFP / RFI requirement compliance triage | legal_compliance | 21 | 0 |
| BIZ-3 | Known-incident linking | infra_ops | 10 | 0 |
| BIZ-30 | Quote / discount approval routing | finance_insurance | 34 | 7 |
| BIZ-31 | Website chat visitor intent pre-routing | speech_dialogue | 1,339 | 3 |
| BIZ-32 | Next-best-action after meeting notes | speech_dialogue | 8 | 0 |
| BIZ-33 | Industry / vertical tagging of accounts | business_ops | 6,243 | 10 |
| BIZ-34 | Partner deal-registration conflict check | finance_insurance | 4,038 | 1 |
| BIZ-35 | Outreach compliance pre-send gate | legal_compliance | 11 | 0 |
| BIZ-36 | Brand-safety / suitability classification of placements | commerce_marketing | 28 | 1 |
| BIZ-37 | Brand social-comment moderation and response routing | security_safety | 35,905 | 26 |
| BIZ-38 | Creator / influencer vetting against brand policy | commerce_marketing | 10 | 4 |
| BIZ-39 | Search-query intent and content-mapping at keyword-list scale | language_nlp | 970 | 4 |
| BIZ-4 | Refund / credit eligibility pre-decision | finance_insurance | 20,853 | 5 |
| BIZ-40 | Content-brief conformance QA | commerce_marketing | 11 | 0 |
| BIZ-41 | Cannibalization and thin-content pair judgments | commerce_marketing | 404 | 6 |
| BIZ-42 | Ad-platform policy pre-check | commerce_marketing | 3,284 | 4 |
| BIZ-43 | Email-campaign pre-send QA | commerce_marketing | 5 | 6 |
| BIZ-44 | Survey / NPS verbatim coding | commerce_marketing | 11,223 | 5 |
| BIZ-45 | Journalist / outlet fit for PR pitching | commerce_marketing | 35 | 2 |
| BIZ-46 | Affiliate / partner site compliance audit | finance_insurance | 3,274 | 4 |
| BIZ-47 | Real-time personalization segment assignment | hr_workforce | 10 | 0 |
| BIZ-48 | Trend / moment suitability for brand participation | commerce_marketing | 7 | 0 |
| BIZ-49 | DAM asset tagging and rights flags | commerce_marketing | 8,852 | 8 |
| BIZ-5 | Bot-to-human handoff trigger | support_ops | 11,752 | 11 |
| BIZ-50 | Localization QA of marketing copy | commerce_marketing | 110 | 5 |
| BIZ-51 | Brand-inbox request triage (sponsorship, partnership, press) | commerce_marketing | 14 | 0 |
| BIZ-52 | Hierarchical product categorization | commerce_marketing | 7,413 | 1 |
| BIZ-53 | Closed-vocabulary attribute normalization | math_reasoning | 7,191 | 0 |
| BIZ-54 | Listing policy-violation detection | travel_hospitality | 532 | 0 |
| BIZ-55 | Duplicate and variant product matching | commerce_marketing | 1,704 | 5 |
| BIZ-56 | Review authenticity and policy scoring | business_ops | 7,340 | 0 |
| BIZ-57 | Return-request triage and disposition | speech_dialogue | 3,315 | 10 |
| BIZ-58 | Chargeback evidence sufficiency and response routing | finance_insurance | 6 | 0 |
| BIZ-59 | Order-anomaly pre-screen (text signals) | business_ops | 7 | 0 |
| BIZ-6 | Pre-survey CSAT and churn-risk prediction | finance_insurance | 11,752 | 11 |
| BIZ-60 | Seller onboarding document and completeness checks (KYB) | finance_insurance | 4,032 | 1 |
| BIZ-61 | Buyer-seller message monitoring | infra_ops | 33 | 0 |
| BIZ-62 | Product Q&A routing and answerability | support_ops | 376 | 0 |
| BIZ-63 | Search relevance judgments for ranking eval and training | commerce_marketing | 1,695 | 5 |
| BIZ-64 | Listing image and content quality QA | travel_hospitality | 207 | 0 |
| BIZ-65 | Price-plausibility and bundle-reading checks | finance_insurance | 7,191 | 0 |
| BIZ-66 | Promo / coupon abuse judgment | security_safety | 15 | 0 |
| BIZ-67 | Subscription cancellation save-offer selection | finance_insurance | 14 | 0 |
| BIZ-68 | Marketplace dispute adjudication triage (INR / SNAD) | business_ops | 806 | 5 |
| BIZ-69 | Product-safety incident detection in reviews and tickets | support_ops | 23,199 | 19 |
| BIZ-7 | Macro / canned-response selection | business_ops | 14 | 0 |
| BIZ-70 | Job-description compliance and bias pre-check | legal_compliance | 8,015 | 6 |
| BIZ-71 | Candidate reply classification for recruiter inbox | hr_workforce | 5 | 6 |
| BIZ-72 | Interview scorecard consistency and non-job-related-comment flags | hr_workforce | 42 | 1 |
| BIZ-73 | Application fraud and duplicate-applicant signals | security_safety | 6,275 | 2 |
| BIZ-74 | Employee helpdesk routing and policy answerability | hr_workforce | 4,534 | 10 |
| BIZ-75 | Leave and accommodation request classification and documentation completeness | legal_compliance | 10 | 0 |
| BIZ-76 | Exit-interview and engagement-survey theme coding with attrition risk | hr_workforce | 8,666 | 7 |
| BIZ-77 | Timesheet and absence justification review | hr_workforce | 19 | 0 |
| BIZ-78 | Employee-relations case intake triage | hr_workforce | 3 | 0 |
| BIZ-79 | Background-check adjudication against a matrix | hr_workforce | 11 | 0 |
| BIZ-8 | Product-feedback (VoC) hierarchical tagging | commerce_marketing | 24,748 | 17 |
| BIZ-80 | Onboarding document classification and completeness | language_nlp | 26 | 0 |
| BIZ-81 | Skills tagging from job history | business_ops | 6,997 | 12 |
| BIZ-82 | Performance-review calibration flags | science_eng | 12,353 | 3 |
| BIZ-83 | Reference and employment-verification response scoring | business_ops | 3 | 1 |
| BIZ-84 | Three-way-match exception cause classification | code_software | 1,412 | 0 |
| BIZ-85 | Duplicate invoice and payment detection | finance_insurance | 15 | 0 |
| BIZ-86 | GL, cost-center and tax-code coding of non-PO invoices and expenses | finance_insurance | 577 | 0 |
| BIZ-87 | Vendor bank-detail change and BEC risk scoring | finance_insurance | 3,277 | 4 |
| BIZ-88 | Collections reply classification and next action | finance_insurance | 2,508 | 0 |
| BIZ-89 | Cash application matching and short-pay reason coding | business_ops | 1,048 | 0 |
| BIZ-9 | Abusive / threatening contact detection and agent protection | security_safety | 40,791 | 28 |
| BIZ-90 | Expense-report policy compliance | finance_insurance | 14,348 | 13 |
| BIZ-91 | Purchase-requisition classification and review routing | business_ops | 11,013 | 6 |
| BIZ-92 | Contract-rate compliance of invoice lines | finance_insurance | 963 | 0 |
| BIZ-93 | Month-end journal-entry plausibility review | data_quality | 1,550 | 7 |
| BIZ-94 | Audit evidence request (PBC) classification and sufficiency | finance_insurance | 2,186 | 6 |
| BIZ-95 | Tax determinations from text: taxability and 1099 vendor status | finance_insurance | 13 | 0 |
| BIZ-96 | Retailer deduction validity (CPG/distribution) | commerce_marketing | 16 | 11 |
| BIZ-97 | Capex vs opex classification of invoice descriptions | finance_insurance | 24 | 0 |
| BIZ-98 | RFQ bid qualitative compliance scoring | legal_compliance | 561 | 3 |
| BIZ-99 | Spend classification for spend analytics | business_ops | 1,776 | 5 |
| DATA-1 | Source-to-target column mapping (schema matching at feed onboarding) | code_software | 1,824 | 2 |
| DATA-10 | Cross-field plausibility check (row sanity) | data_quality | 6,289 | 0 |
| DATA-100 | Answer-vs-reference equivalence | data_quality | 922 | 5 |
| DATA-101 | Regression classification (old answer vs new answer) | math_reasoning | 328 | 2 |
| DATA-102 | Pre-index PII / secret gate on chunks | security_safety | 15,489 | 15 |
| DATA-103 | Re-index materiality on document update | math_reasoning | 2,593 | 6 |
| DATA-104 | Document sensitivity classification for RAG ACLs | llm_rag_data | 571 | 0 |
| DATA-105 | Failure-taxonomy labeling of bad RAG traces | finance_insurance | 10,021 | 5 |
| DATA-106 | Indirect prompt-injection detection in retrieved documents | llm_rag_data | 820 | 3 |
| DATA-107 | Multi-part question coverage check | finance_insurance | 964 | 2 |
| DATA-108 | Pre-labeling with confidence-gated human review (active-learning queue) | llm_rag_data | 339 | 1 |
| DATA-109 | Label QA and annotator-disagreement adjudication | data_quality | 1,033 | 5 |
| DATA-11 | Canonical value normalization (country, state, unit, status codes) | math_reasoning | 2,011 | 2 |
| DATA-110 | Guideline-clause applicability | legal_compliance | 21,944 | 18 |
| DATA-111 | Instruction-prompt quality scoring for SFT data | llm_rag_data | 396 | 3 |
| DATA-112 | SFT response quality rubric | llm_rag_data | 25,336 | 29 |
| DATA-113 | Preference-pair judgment with position-swap self-consistency | data_quality | 38,821 | 38 |
| DATA-114 | Synthetic-sample novelty vs seed set | llm_rag_data | 914 | 1 |
| DATA-115 | Synthetic-row constraint conformance (NL business rules) | commerce_marketing | 6,289 | 0 |
| DATA-116 | Benchmark contamination check | llm_rag_data | 15 | 0 |
| DATA-117 | Educational / pretraining quality scoring (FineWeb-Edu style) | education | 11,252 | 12 |
| DATA-118 | Toxicity / adult / violence filter for pretraining corpora | security_safety | 555 | 5 |
| DATA-119 | Boilerplate / navigation text detection | llm_rag_data | 5,838 | 4 |
| DATA-12 | Cross-system category crosswalk (taxonomy-to-taxonomy) | finance_insurance | 13,249 | 9 |
| DATA-120 | Machine-generated text detection for corpus hygiene | llm_rag_data | 6,867 | 2 |
| DATA-121 | Domain / mixture classification of pretraining documents | llm_rag_data | 19,195 | 6 |
| DATA-122 | Code-file quality scoring for code corpora | llm_rag_data | 15 | 1 |
| DATA-123 | License detection from headers/README text | legal_compliance | 369 | 5 |
| DATA-124 | ASR transcript quality gate | speech_dialogue | 12 | 1 |
| DATA-125 | Hard-negative verification for embedding training | llm_rag_data | 416 | 1 |
| DATA-126 | Slice / attribute tagging for error analysis | math_reasoning | 12 | 0 |
| DATA-127 | Out-of-distribution / novel-class flag | llm_rag_data | 2,449 | 12 |
| DATA-128 | Reward-model-style rubric scoring for RL | llm_rag_data | 7,836 | 10 |
| DATA-129 | Multi-turn dialogue per-turn quality | speech_dialogue | 22 | 2 |
| DATA-13 | Document type routing at ingestion | llm_rag_data | 5,172 | 0 |
| DATA-130 | Translation quality estimation (reference-free) | math_reasoning | 17,566 | 9 |
| DATA-131 | Summary consistency and coverage check | finance_insurance | 16,153 | 6 |
| DATA-132 | Multi-policy fan-out classification with severity | llm_rag_data | 402 | 0 |
| DATA-133 | Review-queue prioritization score | llm_rag_data | 15 | 0 |
| DATA-134 | Appeal triage | security_safety | 6 | 1 |
| DATA-135 | Thread-context moderation | security_safety | 6,045 | 11 |
| DATA-136 | User-report validity scoring | security_safety | 406 | 3 |
| DATA-137 | Profile, username and bio moderation | security_safety | 20 | 0 |
| DATA-138 | Ad creative and landing-page policy review | security_safety | 18 | 0 |
| DATA-139 | Marketplace prohibited-item detection (taxonomy beam) | finance_insurance | 25 | 0 |
| DATA-14 | OCR / extraction quality gate | language_nlp | 29,822 | 27 |
| DATA-140 | Counterfeit-listing text signals | travel_hospitality | 329 | 2 |
| DATA-141 | Child-safety signal escalation gate (always-escalate) | security_safety | 16 | 0 |
| DATA-142 | Self-harm risk tiering | security_safety | 524 | 5 |
| DATA-143 | Harassment: targeted vs. general | language_nlp | 9 | 1 |
| DATA-144 | Claim checkability and known-debunk matching | finance_insurance | 551 | 1 |
| DATA-145 | Coordinated / templated content detection | llm_rag_data | 13 | 0 |
| DATA-146 | Scam conversation pattern detection (romance, investment, gift card) | security_safety | 14 | 1 |
| DATA-147 | Impersonation detection (brand, public figure, staff) | security_safety | 3,292 | 4 |
| DATA-148 | Doxxing / personal-information exposure | hr_workforce | 690 | 2 |
| DATA-149 | Age-appropriateness rating of content | security_safety | 20 | 6 |
| DATA-15 | Address component role assignment | geo_place | 1,071 | 4 |
| DATA-150 | Moderator-decision QA audit | security_safety | 8,752 | 11 |
| DATA-151 | Policy-version relabeling (policy drift) | llm_rag_data | 122 | 0 |
| DATA-152 | Live chat / stream moderation | security_safety | 16 | 0 |
| DATA-153 | Review-platform moderation (off-topic, incentivized, competitor attack) | security_safety | 7,340 | 0 |
| DATA-154 | Extremist-content taxonomy after hash-match miss | finance_insurance | 12 | 0 |
| DATA-155 | AML alert level-1 triage | finance_insurance | 13 | 0 |
| DATA-156 | Sanctions / watchlist name-hit adjudication | finance_insurance | 7,386 | 5 |
| DATA-157 | KYC document-vs-application consistency | finance_insurance | 16 | 4 |
| DATA-158 | KYB business plausibility and industry classification | finance_insurance | 31 | 0 |
| DATA-159 | Chargeback reason coding and evidence sufficiency | finance_insurance | 22 | 0 |
| DATA-16 | Free-text -> typed feature columns (map-reduce featurization) | data_quality | 8,212 | 8 |
| DATA-160 | Account-takeover signals in support contacts | security_safety | 54 | 0 |
| DATA-161 | Promo / referral abuse linked-account judgment | security_safety | 15 | 1 |
| DATA-162 | Refund / return claim abuse text screening | security_safety | 23 | 0 |
| DATA-163 | Business email compromise / invoice-fraud detection | security_safety | 9 | 1 |
| DATA-164 | Fake-signup detection from text signals | security_safety | 17 | 0 |
| DATA-165 | Ad-fraud / made-for-advertising site quality | security_safety | 24 | 0 |
| DATA-166 | Bot / inauthentic engagement detection | llm_rag_data | 10 | 0 |
| DATA-167 | Payment memo / reference risk screening | finance_insurance | 16 | 0 |
| DATA-168 | Seller onboarding risk classification | ux_product | 26 | 0 |
| DATA-169 | Expense receipt policy compliance | finance_insurance | 3,498 | 3 |
| DATA-17 | Data-retention / governance class assignment | legal_compliance | 20 | 1 |
| DATA-170 | Insurance / warranty claim narrative triage | finance_insurance | 13,142 | 4 |
| DATA-171 | SOC alert triage (true positive likelihood) | infra_ops | 1 | 0 |
| DATA-172 | Vendor security questionnaire answer scoring | finance_insurance | 23 | 0 |
| DATA-173 | Transaction narrative categorization for risk features | finance_insurance | 401 | 4 |
| DATA-174 | Job-posting fraud detection | security_safety | 15 | 0 |
| DATA-175 | Loan / credit application narrative consistency | finance_insurance | 16 | 0 |
| DATA-176 | Input guardrail fan-out (one call, many policies) | security_safety | 15,497 | 17 |
| DATA-177 | Output guardrail fan-out | security_safety | 41,827 | 32 |
| DATA-178 | Topic / scope adherence for narrow assistants | language_nlp | 419 | 1 |
| DATA-179 | Tool-call intent alignment gate (agent safety) | speech_dialogue | 6,004 | 6 |
| DATA-18 | Data-catalog description validation | data_quality | 699 | 1 |
| DATA-180 | Destructive / irreversible action gate | security_safety | 14 | 0 |
| DATA-181 | Trajectory success labeling (post-hoc) | llm_rag_data | 1,582 | 10 |
| DATA-182 | Conversation outcome and sentiment features (observability) | infra_ops | 9,098 | 6 |
| DATA-183 | Prompt regression pairwise judge with self-consistency | math_reasoning | 7,006 | 9 |
| DATA-184 | Rubric grading at scale (LLM-judge replacement) | education | 57,952 | 56 |
| DATA-185 | Eval-set hygiene (duplicate, ambiguous, mislabeled items) | llm_rag_data | 339 | 1 |
| DATA-186 | Failure-taxonomy clustering of bad outputs | finance_insurance | 7,362 | 10 |
| DATA-187 | Model routing by query difficulty and risk | code_software | 2,737 | 3 |
| DATA-188 | Semantic-cache validity check | llm_rag_data | 261 | 0 |
| DATA-189 | Streaming partial-output safety check | llm_rag_data | 10 | 1 |
| DATA-19 | Streaming event semantic dedup | data_quality | 23 | 1 |
| DATA-190 | System-prompt leakage detection | security_safety | 7 | 0 |
| DATA-191 | Format compliance check (structural, not arithmetic) | legal_compliance | 7,287 | 16 |
| DATA-192 | Refusal / over-refusal classification of outputs | security_safety | 6,855 | 7 |
| DATA-193 | Red-team attack-success labeling | security_safety | 2,656 | 4 |
| DATA-194 | Data-exfiltration pattern detection in agent outputs | security_safety | 4 | 6 |
| DATA-195 | User-feedback triage (thumbs-down reasons) | llm_rag_data | 201 | 2 |
| DATA-196 | Bias probe pair comparison | security_safety | 1,417 | 6 |
| DATA-197 | Agent step-stagnation / loop detection | llm_rag_data | 17,383 | 15 |
| DATA-198 | Persona and tone consistency scoring | hr_workforce | 11 | 1 |
| DATA-2 | Semantic type inference for untyped columns | data_quality | 3,718 | 12 |
| DATA-20 | Data-observability alert noise classification | infra_ops | 2,828 | 3 |
| DATA-21 | Secret/PII detection in log lines (regex-gated cascade) | security_safety | 15 | 1 |
| DATA-22 | Product matching across retailers/marketplaces | commerce_marketing | 4,008 | 5 |
| DATA-23 | Legal-entity matching with corporate-structure awareness | legal_compliance | 561 | 5 |
| DATA-24 | Person record linkage replacing the clerical-review band | data_quality | 3,453 | 7 |
| DATA-25 | Physical-address equivalence | science_eng | 2,948 | 0 |
| DATA-26 | Semantic near-duplicate confirmation after MinHash/SemDeDup | commerce_marketing | 13,631 | 16 |
| DATA-27 | Golden-record survivorship per attribute | math_reasoning | 1,284 | 3 |
| DATA-28 | Cluster consistency audit (transitive conflicts) | finance_insurance | 1,216 | 2 |
| DATA-29 | Blocking-candidate pre-filter (cheap gate before expensive matcher) | data_quality | 763 | 2 |
| DATA-3 | Locale/format disambiguation for dates and numbers | commerce_marketing | 5,884 | 9 |
| DATA-30 | Author-name disambiguation | data_quality | 404 | 2 |
| DATA-31 | Bibliographic reference to canonical record matching | data_quality | 1,216 | 2 |
| DATA-32 | Job-title normalization to a canonical occupation taxonomy | finance_insurance | 4,196 | 3 |
| DATA-33 | Merchant-string normalization for bank transactions | finance_insurance | 543 | 1 |
| DATA-34 | POI / place deduplication | data_quality | 2,948 | 0 |
| DATA-35 | Media metadata version matching | code_software | 554 | 0 |
| DATA-36 | Duplicate ticket / issue detection | support_ops | 388 | 2 |
| DATA-37 | CRM contact dedup with "moved employer" distinction | data_quality | 16 | 0 |
| DATA-38 | URL / web page canonicalization | data_quality | 56 | 1 |
| DATA-39 | Triple verification against source sentence | llm_rag_data | 4 | 0 |
| DATA-4 | PII / sensitivity tagging of columns and free-text fields | security_safety | 5,777 | 0 |
| DATA-40 | Relation-type selection for an entity pair | llm_rag_data | 4,736 | 8 |
| DATA-41 | Entity typing over an ontology (hierarchical) | language_nlp | 18,229 | 19 |
| DATA-42 | Entity linking (mention -> KB candidate) | llm_rag_data | 758 | 2 |
| DATA-43 | Coreference antecedent choice | language_nlp | 3,674 | 4 |
| DATA-44 | Conflicting-fact resolution | llm_rag_data | 13 | 5 |
| DATA-45 | Ontology / schema alignment | code_software | 587 | 0 |
| DATA-46 | Property-value plausibility (domain/range/format) | travel_hospitality | 3,453 | 7 |
| DATA-47 | Temporal-qualifier assignment for statements | llm_rag_data | 15 | 0 |
| DATA-48 | Alias / synonym merge decision | code_software | 2,534 | 0 |
| DATA-49 | New-concept placement in an existing taxonomy | finance_insurance | 22,410 | 9 |
| DATA-5 | Null-token semantic detection | data_quality | 2,236 | 9 |
| DATA-50 | KG path selection for question answering | llm_rag_data | 4 | 0 |
| DATA-51 | Link-prediction candidate verification | llm_rag_data | 3,453 | 7 |
| DATA-52 | Cross-document event coreference | language_nlp | 19 | 0 |
| DATA-53 | Confidence-gated semantic predicate cascade | llm_rag_data | 17,005 | 8 |
| DATA-54 | Semantic JOIN predicate | llm_rag_data | 665 | 1 |
| DATA-55 | Semantic GROUP BY bucketing | llm_rag_data | 1,458 | 4 |
| DATA-56 | Semantic ORDER BY / top-k via rubric | llm_rag_data | 451 | 1 |
| DATA-57 | NL-to-SQL candidate verification and selection | code_software | 12,123 | 6 |
| DATA-58 | Metric / dimension resolution in a semantic layer | llm_rag_data | 17,005 | 8 |
| DATA-59 | Analytics question intent routing | language_nlp | 17,005 | 8 |
| DATA-6 | Junk-row detection in semi-structured files | data_quality | 1,730 | 0 |
| DATA-60 | Metric-anomaly explanation routing | math_reasoning | 16 | 1 |
| DATA-61 | Query risk classification for DDL/DML review | llm_rag_data | 6,545 | 3 |
| DATA-62 | Row-level access-policy match (ABAC predicate) | llm_rag_data | 0 | 0 |
| DATA-63 | Open-text survey / interview coding to a codebook | hr_workforce | 32 | 0 |
| DATA-64 | JSON / record conformance to a natural-language spec | commerce_marketing | 23,039 | 10 |
| DATA-65 | Load-time normalization to dimension keys | commerce_marketing | 4,345 | 6 |
| DATA-66 | Slow-query pattern classification (DB ops) | llm_rag_data | 3,164 | 2 |
| DATA-67 | Data-sharing eligibility per row (contract clauses) | legal_compliance | 67 | 5 |
| DATA-68 | Warehouse query purpose attribution | math_reasoning | 20 | 4 |
| DATA-69 | Query intent classification for search routing | language_nlp | 1,695 | 5 |
| DATA-7 | Pipeline failure triage (Airflow/Dagster on_failure) | infra_ops | 2,828 | 3 |
| DATA-70 | Graded relevance judgments for offline evaluation (TREC-style) | education | 4,872 | 5 |
| DATA-71 | Constraint-aware rerank for commerce search | llm_rag_data | 1,467 | 5 |
| DATA-72 | Query-rewrite candidate selection | llm_rag_data | 11 | 0 |
| DATA-73 | Retrieval sufficiency gate (agentic retrieval loop) | llm_rag_data | 13,955 | 6 |
| DATA-74 | Query routing across indexes/collections | finance_insurance | 5 | 1 |
| DATA-75 | Lexical-vs-semantic weighting per query | language_nlp | 13 | 1 |
| DATA-76 | Chunk self-containment scoring (chunking strategy tuning) | llm_rag_data | 413 | 1 |
| DATA-77 | Chunk metadata tagging at index time | llm_rag_data | 24 | 4 |
| DATA-78 | Retrieved-passage dedup before context assembly | data_quality | 1,992 | 4 |
| DATA-79 | Snippet / highlight selection | llm_rag_data | 11,361 | 15 |
| DATA-8 | Schema-change impact classification | code_software | 3,336 | 8 |
| DATA-80 | Query-to-product-category prediction (taxonomy beam) | finance_insurance | 5 | 1 |
| DATA-81 | Autocomplete suggestion filtering | llm_rag_data | 414 | 1 |
| DATA-82 | Zero-result / low-result diagnosis | medicine_health | 10 | 6 |
| DATA-83 | Recency-need detection | llm_rag_data | 16 | 0 |
| DATA-84 | Document staleness relative to the query | llm_rag_data | 14 | 0 |
| DATA-85 | Table retrieval judgment | llm_rag_data | 2,692 | 3 |
| DATA-86 | Code-search relevance judgment | commerce_marketing | 14 | 2 |
| DATA-87 | Click-log label denoising | llm_rag_data | 10 | 0 |
| DATA-88 | Result diversity / same-entity detection | llm_rag_data | 1,974 | 4 |
| DATA-89 | Personalized constraint filtering | hr_workforce | 30 | 0 |
| DATA-9 | CDC record-change materiality | math_reasoning | 921 | 12 |
| DATA-90 | Claim-level faithfulness scoring | finance_insurance | 9,927 | 12 |
| DATA-91 | Context precision per chunk | llm_rag_data | 541 | 2 |
| DATA-92 | Context recall per reference statement | llm_rag_data | 8,894 | 7 |
| DATA-93 | Answer relevance scoring | llm_rag_data | 3 | 6 |
| DATA-94 | Runtime hallucination gate with confidence routing | code_software | 10,327 | 8 |
| DATA-95 | Refusal appropriateness (should have refused / should have answered) | security_safety | 7,528 | 8 |
| DATA-96 | Citation assignment (which chunk supports which sentence) | llm_rag_data | 5,223 | 4 |
| DATA-97 | Multi-hop / decomposition need detection | llm_rag_data | 6,235 | 0 |
| DATA-98 | Conversational rewrite intent preservation | speech_dialogue | 1,630 | 3 |
| DATA-99 | Synthetic eval-question quality gate | llm_rag_data | 16 | 0 |
| DEV-1 | Edit-scope and clobber guard | code_software | 991 | 0 |
| DEV-10 | Duplicate tool-call suppression | speech_dialogue | 1,053 | 9 |
| DEV-100 | OAuth/app scope request review | code_software | 531 | 0 |
| DEV-101 | Agent data-exfiltration attempt detection | security_safety | 27 | 3 |
| DEV-102 | Fuzzing-crash triage | infra_ops | 951 | 2 |
| DEV-103 | Secret-usage pattern review | security_safety | 103 | 0 |
| DEV-104 | Sandbox escape/abuse pattern in agent commands | security_safety | 1,871 | 2 |
| DEV-105 | Alert correlation (same-incident detection) | infra_ops | 1,142 | 2 |
| DEV-106 | Alert urgency and notification channel | infra_ops | 274 | 2 |
| DEV-107 | Incident severity classification from first report | infra_ops | 3,620 | 3 |
| DEV-108 | Runbook step selection | infra_ops | 35 | 0 |
| DEV-109 | Probable-cause category | math_reasoning | 9 | 0 |
| DEV-11 | Checkpoint/snapshot moment detection | code_software | 230 | 0 |
| DEV-110 | Culprit-change ranking | code_software | 2 | 0 |
| DEV-111 | Log mislabel and noisy-logger detection | code_software | 385 | 0 |
| DEV-112 | Tail-based trace sampling judgment | code_software | 56 | 0 |
| DEV-113 | Anomaly confirmation | commerce_marketing | 2 | 0 |
| DEV-114 | Stakeholder-update trigger | code_software | 236 | 0 |
| DEV-115 | Incident-channel message classification | infra_ops | 1,131 | 1 |
| DEV-116 | Postmortem action-item classification and priority | infra_ops | 32 | 0 |
| DEV-117 | Ticket-to-team routing | support_ops | 10,585 | 13 |
| DEV-118 | Auto-remediation safety gate | code_software | 39 | 0 |
| DEV-119 | Kubernetes pod failure classification | infra_ops | 66 | 1 |
| DEV-12 | User-message intent classification in a coding session | language_nlp | 16 | 0 |
| DEV-120 | Synthetic-monitor failure classification | infra_ops | 39 | 0 |
| DEV-121 | Support-ticket to incident linking | support_ops | 8 | 1 |
| DEV-122 | Vendor notice impact assessment | finance_insurance | 6,789 | 5 |
| DEV-123 | Root-cause span selection in a trace | code_software | 3 | 0 |
| DEV-124 | Error-group triage (Sentry-style) | code_software | 2,600 | 6 |
| DEV-125 | Crash-report clustering | infra_ops | 1,582 | 1 |
| DEV-126 | Change-request (CAB) classification | code_software | 4,053 | 2 |
| DEV-127 | Chaos-experiment abort decision | science_eng | 50 | 0 |
| DEV-128 | SLO definition quality review | infra_ops | 181 | 0 |
| DEV-129 | Tool-schema pruning per request | code_software | 68,224 | 13 |
| DEV-13 | Clarify-or-proceed gate | code_software | 406 | 1 |
| DEV-130 | Memory write decision | code_software | 780 | 0 |
| DEV-131 | Memory retrieval judgment | llm_rag_data | 461 | 0 |
| DEV-132 | Memory conflict and supersession | code_software | 398 | 0 |
| DEV-133 | Structured-output semantic validation | llm_rag_data | 6,666 | 1 |
| DEV-134 | Tool-argument plausibility check | data_quality | 14,107 | 10 |
| DEV-135 | Scope/charter enforcement for deployed agents | infra_ops | 1,387 | 2 |
| DEV-136 | Specialist-agent handoff | support_ops | 0 | 5 |
| DEV-137 | Failure-mode labeling of agent traces | llm_rag_data | 4,418 | 2 |
| DEV-138 | Pairwise trajectory preference | llm_rag_data | 3 | 6 |
| DEV-139 | LLM-judge audit | finance_insurance | 402 | 6 |
| DEV-14 | Command backgrounding and timeout selection | code_software | 7,722 | 7 |
| DEV-140 | Streaming early-abort | llm_rag_data | 452 | 1 |
| DEV-141 | Speculative tool prefetch | code_software | 543 | 1 |
| DEV-142 | Request priority for shared LLM gateway | llm_rag_data | 35 | 5 |
| DEV-143 | Sub-agent result pass-through decision | llm_rag_data | 7 | 0 |
| DEV-144 | Dialogue-state slot classification | speech_dialogue | 1,150 | 2 |
| DEV-145 | Best-of-N candidate selection | code_software | 560 | 1 |
| DEV-146 | Prompt-variant selection per request | commerce_marketing | 35 | 0 |
| DEV-147 | Retrieval-depth selection | llm_rag_data | 0 | 0 |
| DEV-148 | Semantic response-cache hit | llm_rag_data | 13 | 0 |
| DEV-149 | Human-review sampling for agent outputs | llm_rag_data | 8 | 0 |
| DEV-15 | Final-report claim grounding | finance_insurance | 1,699 | 1 |
| DEV-150 | Time-budget breach prediction | finance_insurance | 406 | 0 |
| DEV-151 | Symbol-definition disambiguation | data_quality | 328 | 1 |
| DEV-152 | Repository-map pruning for agent context | llm_rag_data | 13 | 0 |
| DEV-153 | Documentation snippet reranking (retrieve-then-judge) | llm_rag_data | 2,038 | 5 |
| DEV-154 | Issue duplicate detection | code_software | 187 | 3 |
| DEV-155 | Bug-report to code-area localization | code_software | 19,617 | 16 |
| DEV-156 | Stack-trace frame-of-interest selection | code_software | 1,636 | 2 |
| DEV-157 | Cross-repo impact of an API change | code_software | 343 | 1 |
| DEV-158 | Answer-source routing for developer questions | code_software | 30 | 0 |
| DEV-159 | Regression-introducing-commit prior for bisect | math_reasoning | 29 | 7 |
| DEV-16 | Prompt-injection detection in tool results | llm_rag_data | 13,885 | 5 |
| DEV-160 | Test-to-code mapping without instrumentation | code_software | 660 | 0 |
| DEV-161 | Example-code quality for retrieval indexes | code_software | 6 | 5 |
| DEV-162 | Code-search query intent classification | language_nlp | 1 | 0 |
| DEV-163 | Inline-completion trigger gating | science_eng | 9 | 0 |
| DEV-164 | Completion-candidate acceptance ranking | code_software | 8 | 0 |
| DEV-165 | Diagnostic quick-fix selection | medicine_health | 572 | 6 |
| DEV-166 | Interruption-worthiness of editor notifications | code_software | 31 | 0 |
| DEV-167 | Terminal error to action-class mapping | code_software | 499 | 6 |
| DEV-168 | Rename-refactor occurrence classification | code_software | 12 | 0 |
| DEV-169 | TODO/FIXME triage | code_software | 561 | 0 |
| DEV-17 | Tool-result trust tiering | code_software | 13,885 | 5 |
| DEV-170 | Clipboard-paste content classification | code_software | 291 | 5 |
| DEV-171 | Local-environment mismatch diagnosis | medicine_health | 1,342 | 10 |
| DEV-172 | Auto-import candidate selection | code_software | 616 | 0 |
| DEV-173 | AI-suggestion display-mode selection | code_software | 7 | 0 |
| DEV-174 | API contract change compatibility (OpenAPI/protobuf/GraphQL) | legal_compliance | 1,344 | 8 |
| DEV-175 | Load-shedding priority at the gateway | code_software | 11 | 6 |
| DEV-176 | Cloud cost anomaly triage | infra_ops | 2 | 0 |
| DEV-177 | Rightsizing recommendation safety | commerce_marketing | 44 | 2 |
| DEV-178 | Just-in-time access request adjudication | code_software | 290 | 0 |
| DEV-179 | Dead-letter-queue message triage | language_nlp | 33 | 4 |
| DEV-18 | Worktree/isolation decision per task | code_software | 410 | 1 |
| DEV-180 | Data-pipeline failure classification | infra_ops | 34 | 5 |
| DEV-181 | Slow-query triage | code_software | 535 | 0 |
| DEV-182 | Internal platform request routing | code_software | 284 | 0 |
| DEV-183 | Cloud configuration drift classification | code_software | 1,528 | 3 |
| DEV-184 | Certificate/domain change request review | code_software | 60 | 0 |
| DEV-185 | Service-catalog entry quality scoring | code_software | 3,321 | 4 |
| DEV-186 | Webhook/event routing under ambiguous event types | infra_ops | 85 | 3 |
| DEV-187 | Multi-tenant abuse pattern classification | security_safety | 2 | 6 |
| DEV-19 | Merge-conflict hunk resolution routing | code_software | 42 | 0 |
| DEV-2 | Sandbox profile selection per shell command | data_quality | 7,716 | 7 |
| DEV-20 | Autonomy level per task | science_eng | 16 | 0 |
| DEV-21 | Session-outcome labeling | code_software | 970 | 0 |
| DEV-22 | Human-takeover prediction | code_software | 145 | 0 |
| DEV-23 | Cost-budget continuation gate | finance_insurance | 132 | 0 |
| DEV-24 | Inter-agent message triage | language_nlp | 608 | 3 |
| DEV-25 | Network-egress domain allow decision | code_software | 8,622 | 7 |
| DEV-26 | Environment bootstrap step selection | code_software | 1,900 | 2 |
| DEV-27 | Commit granularity decision | code_software | 2,133 | 0 |
| DEV-28 | PR-readiness and draft decision | code_software | 12 | 0 |
| DEV-29 | Screenshot/UI verification necessity | code_software | 18 | 0 |
| DEV-3 | Task-completion detection | code_software | 2,065 | 1 |
| DEV-30 | Hook-output routing (PostToolUse) | code_software | 39 | 3 |
| DEV-31 | Review-comment actionability triage | code_software | 4,422 | 4 |
| DEV-32 | Reviewer assignment | code_software | 371 | 1 |
| DEV-33 | Review-depth routing | code_software | 548 | 3 |
| DEV-34 | Dependency-bump auto-merge | code_software | 1,749 | 3 |
| DEV-35 | Stale-PR nudge action | code_software | 21 | 0 |
| DEV-36 | Test-adequacy scoring | code_software | 2,364 | 4 |
| DEV-37 | Comment/docstring staleness after a diff | code_software | 816 | 1 |
| DEV-38 | Misleading-identifier lint | code_software | 390 | 0 |
| DEV-39 | Error-handling quality classification | code_software | 721 | 0 |
| DEV-4 | Plan-adherence drift check | code_software | 99 | 0 |
| DEV-40 | Logging-statement review | code_software | 483 | 1 |
| DEV-41 | Duplicate-implementation detection | code_software | 433 | 3 |
| DEV-42 | Feature-flag cleanup readiness | code_software | 31 | 8 |
| DEV-43 | Database migration safety classification | code_software | 9 | 0 |
| DEV-44 | Public API change and version-bump classification | code_software | 3,845 | 4 |
| DEV-45 | Changelog category classification | code_software | 6,191 | 6 |
| DEV-46 | Label auto-tagging from a large label taxonomy | finance_insurance | 985 | 1 |
| DEV-47 | Review-bot comment suppression | code_software | 423 | 0 |
| DEV-48 | Concurrency-hazard flag | code_software | 106 | 0 |
| DEV-49 | Performance-regression suspicion | math_reasoning | 13 | 0 |
| DEV-5 | Delegate-vs-inline decision | science_eng | 5 | 0 |
| DEV-50 | UI diff accessibility/i18n gap check | code_software | 32 | 0 |
| DEV-51 | License-risk snippet flag | legal_compliance | 1,702 | 4 |
| DEV-52 | Composite code-health scoring | medicine_health | 31 | 1 |
| DEV-53 | Generated/boilerplate file classification | llm_rag_data | 1,203 | 0 |
| DEV-54 | Refactor-safety check (behavior preservation) | code_software | 54 | 0 |
| DEV-55 | Config-file diff risk classification | code_software | 231 | 0 |
| DEV-56 | Predictive test selection | code_software | 4,677 | 2 |
| DEV-57 | Flaky-vs-genuine failure classification | code_software | 435 | 1 |
| DEV-58 | Build-log failure root-cause category | code_software | 1,827 | 1 |
| DEV-59 | CI rerun-worthiness | infra_ops | 5 | 0 |
| DEV-6 | Tool-call dependency detection for batching | code_software | 1,299 | 0 |
| DEV-60 | Pipeline stage skipping | infra_ops | 361 | 4 |
| DEV-61 | Change-freeze exception triage | code_software | 470 | 3 |
| DEV-62 | Canary analysis verdict | infra_ops | 2 | 0 |
| DEV-63 | Deploy-attribution rollback decision | math_reasoning | 36 | 1 |
| DEV-64 | Feature-flag rollout step decision | infra_ops | 28 | 0 |
| DEV-65 | Artifact promotion gate | code_software | 502 | 1 |
| DEV-66 | Build-cache reuse safety | code_software | 170 | 0 |
| DEV-67 | Test quarantine and de-quarantine | code_software | 15 | 0 |
| DEV-68 | Failing-test ownership routing | code_software | 1,469 | 1 |
| DEV-69 | CI runner class selection | infra_ops | 19 | 0 |
| DEV-7 | Post-edit verification action selection | code_software | 1,079 | 0 |
| DEV-70 | Monorepo affected-project detection | code_software | 2,847 | 2 |
| DEV-71 | Release-blocking bug triage | code_software | 1,999 | 2 |
| DEV-72 | Post-deploy smoke-test verdict | infra_ops | 791 | 0 |
| DEV-73 | Terraform/IaC plan resource-action review | code_software | 4,816 | 10 |
| DEV-74 | Kubernetes manifest change risk | infra_ops | 128 | 0 |
| DEV-75 | CI workflow security review | infra_ops | 97 | 0 |
| DEV-76 | Build-time/binary-size regression attribution | math_reasoning | 135 | 2 |
| DEV-77 | Release-readiness checklist auto-fill | code_software | 24 | 0 |
| DEV-78 | Nightly-failure deduplication | data_quality | 19 | 1 |
| DEV-79 | Preview-environment necessity | code_software | 11 | 0 |
| DEV-8 | Tool-result error classification | code_software | 2,467 | 1 |
| DEV-80 | Merge-queue batch composition | code_software | 48 | 9 |
| DEV-81 | SAST finding true/false-positive triage | code_software | 326 | 4 |
| DEV-82 | Vulnerability reachability/exploitability judgment | security_safety | 3,012 | 8 |
| DEV-83 | Leaked-credential validity and urgency | security_safety | 49 | 2 |
| DEV-84 | Dependency-update risk classification | code_software | 2,083 | 17 |
| DEV-85 | Maintainer/publish anomaly classification | infra_ops | 15 | 6 |
| DEV-86 | Security bug-report triage | code_software | 13,991 | 9 |
| DEV-87 | Missing-authorization-check detection | code_software | 114 | 0 |
| DEV-88 | IAM policy statement review | code_software | 48 | 1 |
| DEV-89 | Dockerfile/container hardening check | code_software | 53 | 0 |
| DEV-9 | Retry-vs-replan after failure | code_software | 413 | 0 |
| DEV-90 | Edge/WAF request classification | code_software | 566 | 6 |
| DEV-91 | SIEM/cloud-detection alert triage | infra_ops | 1,159 | 1 |
| DEV-92 | Malicious comment/link detection in issue trackers | code_software | 9 | 0 |
| DEV-93 | Telemetry field PII classification | security_safety | 1,072 | 2 |
| DEV-94 | SBOM license classification | legal_compliance | 1,702 | 4 |
| DEV-95 | Threat-model trigger | security_safety | 408 | 0 |
| DEV-96 | Cryptography-misuse flag | code_software | 70 | 2 |
| DEV-97 | Security-exception request adjudication | code_software | 6,256 | 5 |
| DEV-98 | Cross-tool finding deduplication (SAST/DAST/SCA/pentest) | data_quality | 2,643 | 1 |
| DEV-99 | Runtime security event classification | code_software | 2,433 | 1 |
| PHYS-1 | Step verdict: done / continue / stuck | science_eng | 23,259 | 7 |
| PHYS-10 | Viewport navigation to reveal the target | science_eng | 7 | 0 |
| PHYS-100 | Lab-scheduler step arbitration | science_eng | 18 | 2 |
| PHYS-101 | Hypothesis prioritization | science_eng | 402 | 2 |
| PHYS-102 | Sequencing/omics QC gate | science_eng | 2,400 | 9 |
| PHYS-103 | Editorial desk triage | ux_product | 10,933 | 18 |
| PHYS-104 | Autonomous microscopy next-probe location | science_eng | 2 | 0 |
| PHYS-105 | Simulation-sweep early stopping | science_eng | 249 | 0 |
| PHYS-106 | Beamline/instrument anomaly response | science_eng | 17 | 0 |
| PHYS-107 | Grant and paper routing | finance_insurance | 7,001 | 4 |
| PHYS-108 | Statistical-reporting plausibility flag | math_reasoning | 38 | 0 |
| PHYS-109 | Virtual-screening hit triage | science_eng | 25 | 0 |
| PHYS-11 | Tab and window management | science_eng | 15,656 | 7 |
| PHYS-110 | Synthesis-route step feasibility | science_eng | 7,229 | 4 |
| PHYS-111 | Reaction-condition selection | science_eng | 3,706 | 11 |
| PHYS-112 | HTS well-level artifact flagging | code_software | 6 | 0 |
| PHYS-113 | Crystallization outcome classification | science_eng | 12 | 0 |
| PHYS-114 | Assay-cascade progression | llm_rag_data | 4 | 0 |
| PHYS-115 | Protein-variant library prioritization | commerce_marketing | 10 | 0 |
| PHYS-116 | Cell-culture passage and feed decisions | llm_rag_data | 184 | 0 |
| PHYS-117 | Formulation stability early kill | science_eng | 20 | 1 |
| PHYS-118 | Fed-batch bioprocess feed decision | science_eng | 16 | 0 |
| PHYS-119 | Materials-recipe triage for synthesis | science_eng | 3,085 | 6 |
| PHYS-12 | Error-page, rate-limit and bot-check recognition | science_eng | 12,361 | 3 |
| PHYS-120 | Quote width and skew selection for market making | finance_insurance | 4,033 | 2 |
| PHYS-121 | Cancel/replace vs hold under latency | infra_ops | 17 | 0 |
| PHYS-122 | Execution-algo child-order tactics | science_eng | 15 | 1 |
| PHYS-123 | Venue routing | science_eng | 15 | 0 |
| PHYS-124 | Toxic-flow detection | security_safety | 25 | 2 |
| PHYS-125 | Strategy kill-switch | science_eng | 9,167 | 0 |
| PHYS-126 | Event-driven exposure change | science_eng | 8 | 0 |
| PHYS-127 | Pairs / stat-arb entry and exit | science_eng | 3,204 | 0 |
| PHYS-128 | Options market-making skew adjust | science_eng | 880 | 0 |
| PHYS-129 | Perps/DeFi margin management | science_eng | 3,058 | 2 |
| PHYS-13 | Extraction targeting: which rows and columns to pull | science_eng | 27 | 0 |
| PHYS-130 | Trade-surveillance alert triage | infra_ops | 23 | 1 |
| PHYS-131 | Rebalance trigger | science_eng | 6,931 | 3 |
| PHYS-132 | Prediction-market / sportsbook line adjustment | games_puzzles | 12,445 | 4 |
| PHYS-133 | Arbitrage opportunity acceptance | llm_rag_data | 16 | 0 |
| PHYS-134 | Battery-storage dispatch | science_eng | 519 | 0 |
| PHYS-135 | Demand-response participation | science_eng | 6 | 0 |
| PHYS-136 | FLISR switch selection | science_eng | 51 | 0 |
| PHYS-137 | Renewable curtailment decision | science_eng | 12,936 | 3 |
| PHYS-138 | EV charging load management | science_eng | 1,477 | 12 |
| PHYS-139 | Substation inspection finding disposition | science_eng | 7,741 | 4 |
| PHYS-14 | Cross-site offer selection | science_eng | 17 | 0 |
| PHYS-140 | Wildfire PSPS sector decision | science_eng | 31 | 0 |
| PHYS-141 | Well and pipeline control actions | infra_ops | 6 | 0 |
| PHYS-142 | Home energy management | science_eng | 492 | 0 |
| PHYS-143 | Frequency-response bid band | science_eng | 3,671 | 2 |
| PHYS-144 | Irrigation zone scheduling | math_reasoning | 3,159 | 10 |
| PHYS-145 | Spot-spray trigger | science_eng | 3,558 | 0 |
| PHYS-146 | Harvest-timing readiness | science_eng | 2,410 | 11 |
| PHYS-147 | Livestock health alert | medicine_health | 33 | 1 |
| PHYS-148 | Greenhouse climate setpoint selection | science_eng | 37 | 0 |
| PHYS-149 | Scouting prioritization | science_eng | 934 | 0 |
| PHYS-15 | Assistive-tech intent grounding | language_nlp | 15,827 | 7 |
| PHYS-150 | Grain-storage aeration control | llm_rag_data | 1,716 | 0 |
| PHYS-151 | Fruit-picking target selection | science_eng | 234 | 6 |
| PHYS-152 | Alarm-correlation root cause | llm_rag_data | 203 | 2 |
| PHYS-153 | SON self-healing action | science_eng | 36 | 0 |
| PHYS-154 | Ticket routing | support_ops | 10,765 | 9 |
| PHYS-155 | Change-window go/no-go | science_eng | 11 | 0 |
| PHYS-156 | DDoS mitigation activation | science_eng | 16 | 0 |
| PHYS-157 | Traffic-engineering path selection | science_eng | 42 | 0 |
| PHYS-158 | Field-dispatch prioritization | science_eng | 1,385 | 6 |
| PHYS-159 | Customer-premise troubleshooting step | science_eng | 9 | 11 |
| PHYS-16 | UI regression diff triage | math_reasoning | 10 | 0 |
| PHYS-160 | Network-slice allocation | science_eng | 65 | 0 |
| PHYS-161 | Cell energy-saving mode | science_eng | 53 | 0 |
| PHYS-162 | Map-change triage from probe data | geo_place | 2,138 | 4 |
| PHYS-163 | Entity conflation | science_eng | 51 | 0 |
| PHYS-164 | Imagery tile prioritization for annotation | data_quality | 45 | 0 |
| PHYS-165 | Geocoding candidate selection | geo_place | 567 | 5 |
| PHYS-166 | Feature QA flag | knowledge_qa | 3,728 | 6 |
| PHYS-167 | Damage-assessment triage | science_eng | 58 | 1 |
| PHYS-168 | Route-hazard verification | science_eng | 1,361 | 6 |
| PHYS-169 | Land-use classification from indices | science_eng | 52 | 1 |
| PHYS-17 | Self-healing locator selection | science_eng | 6,085 | 2 |
| PHYS-170 | Dispatch assignment | science_eng | 55 | 0 |
| PHYS-171 | Disruption re-route | science_eng | 15 | 0 |
| PHYS-172 | Spot-load acceptance | science_eng | 24 | 0 |
| PHYS-173 | Next-pick selection in warehouses | science_eng | 48 | 0 |
| PHYS-174 | Crane and yard job sequencing | science_eng | 17 | 1 |
| PHYS-175 | Transit signal priority | science_eng | 1,780 | 8 |
| PHYS-176 | Delivery-exception handling | code_software | 64 | 1 |
| PHYS-177 | Elevator/AGV group dispatch | science_eng | 8 | 0 |
| PHYS-178 | Irregular-operations recovery choice | science_eng | 46 | 0 |
| PHYS-179 | Parcel misroute detection | science_eng | 55 | 0 |
| PHYS-18 | Fast critic for a slow LLM agent | llm_rag_data | 1,584 | 4 |
| PHYS-180 | Production NPC behaviour arbitration | science_eng | 17 | 0 |
| PHYS-181 | Matchmaking and host selection | science_eng | 34 | 0 |
| PHYS-182 | Bot/cheat detection | science_eng | 9 | 0 |
| PHYS-183 | Dynamic difficulty adjustment | code_software | 5,125 | 2 |
| PHYS-184 | Procedural-content curation | science_eng | 4,651 | 2 |
| PHYS-185 | Playtest bot stuck-state detection | science_eng | 34 | 2 |
| PHYS-186 | Card-game play selection | games_puzzles | 3,939 | 13 |
| PHYS-187 | Auction/negotiation sim agents | llm_rag_data | 6 | 0 |
| PHYS-188 | Agent-based simulation micro-decisions | llm_rag_data | 19 | 5 |
| PHYS-189 | Search-guiding move prior | science_eng | 3,905 | 13 |
| PHYS-19 | Document line-role classification from OCR | language_nlp | 9 | 0 |
| PHYS-190 | Screenshot-native computer-use decider | science_eng | 15,827 | 7 |
| PHYS-191 | Long-horizon browser tasks with persistent state | science_eng | 33 | 0 |
| PHYS-192 | Composite control outputs | science_eng | 1,833 | 1 |
| PHYS-193 | Plan-verify micro-reasoning | science_eng | 17,140 | 10 |
| PHYS-194 | Native 10k-way selection | science_eng | 5,968 | 2 |
| PHYS-195 | Multi-modal sensor fusion decider | science_eng | 173 | 1 |
| PHYS-196 | Structured spatial outputs | science_eng | 7,227 | 3 |
| PHYS-197 | Date/arithmetic-aware scheduling decisions | math_reasoning | 24,556 | 23 |
| PHYS-198 | Per-frame video decisions | science_eng | 18 | 0 |
| PHYS-199 | Simulator-in-the-loop counterfactual gating | math_reasoning | 24 | 2 |
| PHYS-2 | WebMCP tool vs raw-DOM action routing | science_eng | 175 | 0 |
| PHYS-20 | Desktop screen-state classification | science_eng | 2,893 | 2 |
| PHYS-200 | Language-conditioned embodied policies | science_eng | 17,117 | 10 |
| PHYS-21 | Green-screen (3270/5250) navigation | science_eng | 31 | 0 |
| PHYS-22 | RPA exception routing | code_software | 1,328 | 6 |
| PHYS-23 | Mobile permission and onboarding dialog policy | ux_product | 28 | 0 |
| PHYS-24 | On-device notification triage | ux_product | 36 | 0 |
| PHYS-25 | Email-to-ERP posting decisions | science_eng | 542 | 0 |
| PHYS-26 | Spreadsheet column-role and row-cleaning actions | science_eng | 479 | 1 |
| PHYS-27 | Back-office workflow next-step selection | science_eng | 10,716 | 9 |
| PHYS-28 | Pre-submit consistency gate | data_quality | 9,314 | 6 |
| PHYS-29 | App hang and crash recovery | infra_ops | 21 | 1 |
| PHYS-3 | Form-field to profile-field mapping | data_quality | 7,826 | 3 |
| PHYS-30 | Exploratory UI crawler for mobile QA | knowledge_qa | 8 | 4 |
| PHYS-31 | Grasp candidate selection | science_eng | 419 | 0 |
| PHYS-32 | Pick-order selection in cluttered bins | science_eng | 18 | 0 |
| PHYS-33 | Behavior-tree node arbitration for manipulation | science_eng | 17,117 | 10 |
| PHYS-34 | Navigation recovery behavior selection | science_eng | 15 | 1 |
| PHYS-35 | Social navigation micro-decisions | science_eng | 11 | 0 |
| PHYS-36 | AMR intersection and deadlock arbitration | science_eng | 7 | 1 |
| PHYS-37 | Force-guided insertion phase control | science_eng | 23 | 1 |
| PHYS-38 | Grasp-slip and drop detection | science_eng | 21 | 0 |
| PHYS-39 | Human handover readiness | science_eng | 172 | 1 |
| PHYS-4 | Pagination and infinite-scroll stop decision | ux_product | 16 | 0 |
| PHYS-40 | Multi-robot task allocation | science_eng | 20 | 0 |
| PHYS-41 | Legged-robot gait and mode selection | science_eng | 5 | 0 |
| PHYS-42 | Balance-loss safe-stop for humanoids | science_eng | 20 | 0 |
| PHYS-43 | Shared-autonomy blending | science_eng | 19 | 1 |
| PHYS-44 | Robot-learning episode curation | science_eng | 14 | 0 |
| PHYS-45 | End-effector/tool selection | science_eng | 9,830 | 10 |
| PHYS-46 | Household robot placement target selection | science_eng | 17,117 | 10 |
| PHYS-47 | Lane-change go/no-go | science_eng | 434 | 2 |
| PHYS-48 | Unprotected turn and intersection creep | speech_dialogue | 188 | 0 |
| PHYS-49 | Merge and zipper yielding | code_software | 95 | 0 |
| PHYS-5 | Result-list candidate ranking (search results, listings, catalogs) | travel_hospitality | 553 | 2 |
| PHYS-50 | Pedestrian and cyclist crossing intent | science_eng | 51 | 0 |
| PHYS-51 | Drive-log scenario mining | science_eng | 1,537 | 10 |
| PHYS-52 | Adversarial NPC behaviour in simulation | security_safety | 188 | 0 |
| PHYS-53 | Handover / disengagement request | science_eng | 501 | 8 |
| PHYS-54 | Parking maneuver step selection | science_eng | 5 | 0 |
| PHYS-55 | Driver-monitoring escalation | science_eng | 501 | 8 |
| PHYS-56 | Speed-profile choice in special zones | data_quality | 29 | 0 |
| PHYS-57 | Emergency-vehicle response | science_eng | 20 | 1 |
| PHYS-58 | Remote-assist queue triage | infra_ops | 10,716 | 9 |
| PHYS-59 | Scenario-parameter falsification | science_eng | 561 | 0 |
| PHYS-6 | Overlay, modal and consent-banner handling | legal_compliance | 9 | 0 |
| PHYS-60 | Inspection waypoint prioritization | science_eng | 18 | 0 |
| PHYS-61 | Abort / return-to-home | speech_dialogue | 2,348 | 2 |
| PHYS-62 | Landing-site selection | science_eng | 8 | 0 |
| PHYS-63 | Detect-and-avoid maneuver | science_eng | 7 | 0 |
| PHYS-64 | Delivery drop go/no-go | science_eng | 2,231 | 0 |
| PHYS-65 | Swarm role and target assignment | science_eng | 20 | 0 |
| PHYS-66 | Survey tile re-fly decision | science_eng | 20 | 1 |
| PHYS-67 | Launch / hold weather decision | science_eng | 2,515 | 0 |
| PHYS-68 | Gimbal and capture action | science_eng | 17 | 1 |
| PHYS-69 | Anomaly follow-up on inspection | science_eng | 17,311 | 3 |
| PHYS-7 | Auth-wall and credential-field detection with handoff | security_safety | 12,361 | 3 |
| PHYS-70 | Alarm-flood triage (ISA-18.2) | science_eng | 243 | 0 |
| PHYS-71 | Operator next-action advisor on DCS | science_eng | 6,206 | 9 |
| PHYS-72 | AOI/vision defect disposition | science_eng | 8,751 | 4 |
| PHYS-73 | SPC out-of-control response | science_eng | 4,364 | 0 |
| PHYS-74 | Predictive-maintenance work-order triage | science_eng | 10,570 | 9 |
| PHYS-75 | Job sequencing and changeover selection | science_eng | 14 | 0 |
| PHYS-76 | CNC adaptive feed override | science_eng | 21 | 0 |
| PHYS-77 | Weld quality in-process gate | science_eng | 30 | 1 |
| PHYS-78 | Cobot speed-and-separation mode | science_eng | 13 | 2 |
| PHYS-79 | Packaging-line jam recovery | science_eng | 21 | 1 |
| PHYS-8 | Irreversible-action gate (send, submit, purchase, delete) | security_safety | 15,656 | 7 |
| PHYS-80 | Golden-batch deviation response | science_eng | 6,210 | 9 |
| PHYS-81 | Skip-lot incoming inspection | science_eng | 9,314 | 6 |
| PHYS-82 | Rework station routing | science_eng | 7,956 | 4 |
| PHYS-83 | Root-cause code shortlisting | travel_hospitality | 14,009 | 20 |
| PHYS-84 | Machine idle / standby decision | science_eng | 412 | 0 |
| PHYS-85 | Edge upload gating | science_eng | 20,840 | 4 |
| PHYS-86 | Sensor fault vs real anomaly | science_eng | 20,840 | 4 |
| PHYS-87 | Adaptive sampling rate | science_eng | 20,840 | 4 |
| PHYS-88 | Occupancy-driven HVAC and lighting actions | science_eng | 16 | 0 |
| PHYS-89 | Access and badge anomaly | science_eng | 18 | 0 |
| PHYS-9 | Duplicate-element disambiguation | data_quality | 26,663 | 10 |
| PHYS-90 | Cold-chain excursion response | medicine_health | 3,989 | 9 |
| PHYS-91 | Wearable fall confirmation | science_eng | 10 | 0 |
| PHYS-92 | Acoustic/camera-trap trigger | science_eng | 9 | 0 |
| PHYS-93 | Water-network leak flagging | science_eng | 873 | 2 |
| PHYS-94 | Firmware canary promotion | infra_ops | 29 | 0 |
| PHYS-95 | Sensor-fusion trust arbitration | science_eng | 22 | 0 |
| PHYS-96 | Systematic-review abstract screening | science_eng | 9,476 | 12 |
| PHYS-97 | Evidence-table claim relevance | finance_insurance | 1,181 | 1 |
| PHYS-98 | Next-experiment selection from a BO batch | science_eng | 3,294 | 11 |
| PHYS-99 | Experiment run QC triage | science_eng | 3,476 | 1 |
| UX-1 | Intent-ahead prefetch and view warming | language_nlp | 299 | 2 |
| UX-10 | Checkout abandonment prediction and friction diagnosis | medicine_health | 411 | 1 |
| UX-100 | Autocorrect suppression for intentional words | language_nlp | 8 | 0 |
| UX-101 | Candidate ranking for next-word and emoji strip | language_nlp | 6 | 0 |
| UX-102 | Language switch detection while typing | language_nlp | 5 | 4 |
| UX-103 | Sentence-end and capitalization prediction | language_nlp | 555 | 1 |
| UX-104 | Swipe-typing candidate disambiguation | data_quality | 6 | 0 |
| UX-105 | Dictation command vs. content | speech_dialogue | 16 | 2 |
| UX-106 | Field-type inference for keyboard and autofill | ux_product | 906 | 5 |
| UX-107 | Paste-transform selection | ux_product | 20,701 | 7 |
| UX-108 | Stylus stroke intent | language_nlp | 545 | 6 |
| UX-109 | Stuck-typing help trigger | ux_product | 20 | 0 |
| UX-11 | Upsell and paywall moment gating | commerce_marketing | 9 | 0 |
| UX-110 | Best-frame selection in bursts | ux_product | 10 | 0 |
| UX-111 | Screenshot purpose routing | ux_product | 402 | 1 |
| UX-112 | Library junk and near-duplicate detection | code_software | 21 | 0 |
| UX-113 | Memory-worthiness scoring | ux_product | 405 | 2 |
| UX-114 | Share-safety scan | ux_product | 10 | 0 |
| UX-115 | Camera mode suggestion | ux_product | 4 | 0 |
| UX-116 | Document capture readiness | ux_product | 4 | 0 |
| UX-117 | Captured document routing | ux_product | 5 | 0 |
| UX-118 | Photo search re-ranking | ux_product | 17 | 0 |
| UX-119 | Video highlight scoring | llm_rag_data | 9,477 | 4 |
| UX-12 | Smart default pre-selection in pickers | ux_product | 416 | 1 |
| UX-120 | Household state inference | ux_product | 15,396 | 14 |
| UX-121 | Is-this-normal anomaly judgment | commerce_marketing | 18 | 0 |
| UX-122 | Automation conflict resolution | ux_product | 4,354 | 13 |
| UX-123 | Camera alert worthiness | infra_ops | 6 | 0 |
| UX-124 | Appliance run-window recommendation | commerce_marketing | 6,696 | 4 |
| UX-125 | Comfort pre-conditioning | ux_product | 14 | 0 |
| UX-126 | Multi-room voice target resolution | speech_dialogue | 12,495 | 2 |
| UX-127 | Activity classification on wearables | science_eng | 13 | 0 |
| UX-128 | Wearable notification receptivity | science_eng | 10 | 0 |
| UX-129 | Fall-alert false-alarm second opinion | finance_insurance | 15 | 0 |
| UX-13 | Contextual copy and layout variant selection | commerce_marketing | 394 | 2 |
| UX-130 | Vitals notification worthiness | ux_product | 34 | 0 |
| UX-131 | Appliance error triage | ux_product | 15 | 4 |
| UX-132 | Sound-event escalation for monitors | support_ops | 449 | 5 |
| UX-133 | Doorbell visitor intent | language_nlp | 6 | 4 |
| UX-134 | Transaction categorization at swipe time | finance_insurance | 299 | 0 |
| UX-135 | Merchant normalization | finance_insurance | 4 | 0 |
| UX-136 | Subscription creep and forgotten trials | finance_insurance | 6 | 0 |
| UX-137 | Budget nudge timing | finance_insurance | 21 | 0 |
| UX-138 | Financial document classification | finance_insurance | 15 | 0 |
| UX-139 | Adaptive habit reminder | ux_product | 10 | 6 |
| UX-14 | Arrival-time urgency and reply-need scoring | ux_product | 19,419 | 9 |
| UX-140 | Free-text meal parsing judgment | ux_product | 0 | 0 |
| UX-141 | Symptom note triage (informational) | medicine_health | 4 | 0 |
| UX-142 | Medication adherence inference | ux_product | 11 | 0 |
| UX-143 | Training session adaptation | ux_product | 16 | 0 |
| UX-144 | Goal drift detection | ux_product | 14 | 0 |
| UX-145 | Bill-split item attribution | math_reasoning | 428 | 0 |
| UX-146 | Dynamic difficulty from telemetry | science_eng | 10 | 0 |
| UX-147 | Hint timing | ux_product | 26 | 0 |
| UX-148 | In-game chat toxicity triage | security_safety | 15,706 | 12 |
| UX-149 | Companion NPC tactical choice | science_eng | 5 | 0 |
| UX-15 | Waiting-on state for threads | ux_product | 3,782 | 4 |
| UX-150 | Procedural content curation | science_eng | 603 | 0 |
| UX-151 | Scripted-input detection | ux_product | 21 | 0 |
| UX-152 | Clip-worthiness for auto-highlights | llm_rag_data | 9,864 | 5 |
| UX-153 | Tutorial skip judgment | education | 17 | 2 |
| UX-154 | Player archetype profiling | data_quality | 28 | 0 |
| UX-155 | Branch prediction for interactive fiction | code_software | 405 | 1 |
| UX-156 | Streamer chat command routing | speech_dialogue | 18 | 0 |
| UX-157 | Return-prompt appropriateness for mobile games | speech_dialogue | 17 | 2 |
| UX-158 | Layer grouping and naming suggestion | ux_product | 1,373 | 0 |
| UX-159 | Palette harmony and mood judgment | ux_product | 5,073 | 2 |
| UX-16 | Pre-send guard | ux_product | 12 | 0 |
| UX-160 | Font pairing fit | ux_product | 5,073 | 2 |
| UX-161 | DAW arrangement and clash judgment | ux_product | 6 | 0 |
| UX-162 | Sentence-level edit action judgments | ux_product | 416 | 1 |
| UX-163 | Cut-point suggestion from transcript | speech_dialogue | 9,866 | 5 |
| UX-164 | Design-system deviation lint | code_software | 21 | 0 |
| UX-165 | Tool routing from natural-language intent | language_nlp | 18,878 | 28 |
| UX-166 | Generated-asset curation against a brief | ux_product | 12 | 0 |
| UX-167 | Endpointing from partial transcript | speech_dialogue | 426 | 0 |
| UX-168 | Addressed-to-assistant rejection | geo_place | 9 | 6 |
| UX-169 | Barge-in interpretation | ux_product | 16 | 0 |
| UX-17 | Receipt-time bundling category | ux_product | 13,924 | 9 |
| UX-170 | Clarification gating | ux_product | 22,426 | 19 |
| UX-171 | Local vs. cloud routing | ux_product | 4,354 | 13 |
| UX-172 | Confirmation gating by irreversibility | commerce_marketing | 16,849 | 15 |
| UX-173 | Follow-up expectation | ux_product | 216 | 4 |
| UX-174 | Tone and urgency for response style | ux_product | 15,706 | 12 |
| UX-175 | Tab revisit likelihood | ux_product | 4 | 4 |
| UX-176 | Reader-mode auto-trigger | ux_product | 13 | 0 |
| UX-177 | Cookie-banner button selection | ux_product | 100 | 0 |
| UX-178 | Form-field semantic mapping for autofill | llm_rag_data | 904 | 5 |
| UX-179 | Hover-click prediction for prerendering | ux_product | 2 | 2 |
| UX-18 | Snooze bucket suggestion | ux_product | 94 | 0 |
| UX-180 | Wall detection and alternative routing | ux_product | 21 | 0 |
| UX-181 | Deal legitimacy and dark-pattern flags | ux_product | 4,981 | 5 |
| UX-182 | Session-restore triage | ux_product | 87 | 0 |
| UX-183 | Suspicious-download feel | ux_product | 18 | 0 |
| UX-184 | Search-result re-ranking by intent | language_nlp | 413 | 1 |
| UX-185 | Off-task judgment for focus blockers | ux_product | 1,586 | 2 |
| UX-186 | Review authenticity feel | ux_product | 2,026 | 10 |
| UX-187 | Playback-speed suggestion | ux_product | 9,483 | 4 |
| UX-188 | Next-app prediction and prewarming | ux_product | 288 | 2 |
| UX-189 | Background-refresh worthiness | ux_product | 1,660 | 7 |
| UX-19 | Task-in-email detection | ux_product | 13,924 | 9 |
| UX-190 | Permission-request plausibility | data_quality | 934 | 2 |
| UX-191 | Share-sheet target ranking | ux_product | 5 | 0 |
| UX-192 | Focus-mode auto-switch | ux_product | 453 | 5 |
| UX-193 | Download and file auto-filing | ux_product | 3,960 | 6 |
| UX-194 | Screen-time intervention timing | ux_product | 11 | 0 |
| UX-195 | Widget and live-activity relevance | llm_rag_data | 464 | 2 |
| UX-196 | Cross-device continuation prediction | ux_product | 5 | 0 |
| UX-197 | Selected-text action selection | ux_product | 683 | 4 |
| UX-198 | On-device crash triage | infra_ops | 2,758 | 7 |
| UX-199 | Update timing | ux_product | 388 | 1 |
| UX-2 | Adaptive form skip-logic | math_reasoning | 904 | 5 |
| UX-20 | Sent-mail follow-up worthiness | ux_product | 2,370 | 5 |
| UX-200 | Storage cleanup candidates | llm_rag_data | 18 | 0 |
| UX-21 | Phishing and social-engineering feel | security_safety | 562 | 0 |
| UX-22 | Sender engagement and unsubscribe suggestion | finance_insurance | 11 | 0 |
| UX-23 | Ghost-text acceptance gating for compose | ux_product | 18 | 0 |
| UX-24 | Invite response recommendation | commerce_marketing | 11 | 0 |
| UX-25 | Overlap resolution | ux_product | 7 | 0 |
| UX-26 | Focus-block protection | ux_product | 0 | 0 |
| UX-27 | Event type classification for auto-settings | ux_product | 12 | 0 |
| UX-28 | Running-late detection | ux_product | 20 | 0 |
| UX-29 | Meeting-prep relevance fan-out | speech_dialogue | 425 | 1 |
| UX-3 | Confirmation-dialog gating (undo instead of "Are you sure?") | commerce_marketing | 15,396 | 14 |
| UX-30 | Should-this-be-a-meeting | speech_dialogue | 407 | 1 |
| UX-31 | Reminder context selection | ux_product | 552 | 3 |
| UX-32 | Save-time routing and tagging | ux_product | 26 | 0 |
| UX-33 | Stale-note detection | ux_product | 21 | 0 |
| UX-34 | Backlink suggestion fan-out | llm_rag_data | 455 | 2 |
| UX-35 | Quick-capture type routing | ux_product | 25 | 0 |
| UX-36 | Live line-role formatting | commerce_marketing | 451 | 0 |
| UX-37 | Near-duplicate note detection | ux_product | 75 | 0 |
| UX-38 | Sensitivity gate before sync or share | ux_product | 5 | 3 |
| UX-39 | Column semantics on paste | llm_rag_data | 541 | 1 |
| UX-4 | Onboarding and empty-state tip selection | ux_product | 46 | 0 |
| UX-40 | Row anomaly flagging | ux_product | 8,080 | 8 |
| UX-41 | Fuzzy entity matching for merge and dedupe | code_software | 3,001 | 2 |
| UX-42 | Formula intent from header and neighbors | language_nlp | 8 | 0 |
| UX-43 | Chart-type suggestion | ux_product | 13 | 0 |
| UX-44 | Import structure cleaning | ux_product | 480 | 1 |
| UX-45 | Per-keystroke source routing | ux_product | 2,544 | 4 |
| UX-46 | Clipboard history relevance and secret masking | security_safety | 13 | 0 |
| UX-47 | Window and tab switcher by task | ux_product | 28 | 0 |
| UX-48 | Destructive command disambiguation | security_safety | 8,202 | 14 |
| UX-49 | Resume-likely documents | hr_workforce | 9 | 0 |
| UX-5 | Toast and banner importance sizing | ux_product | 12 | 0 |
| UX-50 | Typo vs. new term | language_nlp | 415 | 1 |
| UX-51 | Completion display gating | ux_product | 12 | 0 |
| UX-52 | Diff hunk risk scoring | code_software | 19,572 | 6 |
| UX-53 | Warning actionability filter | ux_product | 12 | 0 |
| UX-54 | Affected-test selection | code_software | 4,483 | 1 |
| UX-55 | Change-type classification | ux_product | 19,533 | 5 |
| UX-56 | Terminal output triage | ux_product | 7 | 0 |
| UX-57 | Live log severity re-scoring | ux_product | 431 | 0 |
| UX-58 | Dependency bump risk | code_software | 2,501 | 0 |
| UX-59 | Stale notebook cell detection | ux_product | 841 | 0 |
| UX-6 | Error recovery action selection | ux_product | 1,453 | 1 |
| UX-60 | Interrupt-now vs. hold-for-summary | language_nlp | 21 | 2 |
| UX-61 | Notification true-category correction | ux_product | 406 | 1 |
| UX-62 | Burst coalescing | commerce_marketing | 405 | 4 |
| UX-63 | Context-aware delivery mode | ux_product | 407 | 0 |
| UX-64 | Group-chat relevance to me | speech_dialogue | 0 | 0 |
| UX-65 | Quick-reply intent selection | language_nlp | 25 | 2 |
| UX-66 | Conversation heat detection | speech_dialogue | 32,236 | 32 |
| UX-67 | Delay-send regret guard | ux_product | 6 | 0 |
| UX-68 | SMS scam detection | security_safety | 387 | 0 |
| UX-69 | Notification auto-expiry | ux_product | 758 | 4 |
| UX-7 | Progressive disclosure by inferred expertise | legal_compliance | 454 | 0 |
| UX-70 | Muted-thread exception | code_software | 5,563 | 4 |
| UX-71 | Cross-app duplicate suppression | ux_product | 13 | 0 |
| UX-72 | Do-Not-Disturb breakthrough | ux_product | 9 | 0 |
| UX-73 | Response-form selection (react vs. reply) | ux_product | 15,706 | 12 |
| UX-74 | Catch-up summary trigger | language_nlp | 0 | 0 |
| UX-75 | Missed-something detection in playback | ux_product | 9,472 | 4 |
| UX-76 | Cross-platform story deduplication | data_quality | 407 | 1 |
| UX-77 | Next-track fit ranking | ux_product | 456 | 2 |
| UX-78 | Skip prediction and pre-buffering | ux_product | 12 | 5 |
| UX-79 | Chapter boundary and type detection from transcript | speech_dialogue | 9,877 | 5 |
| UX-8 | In-app search intent routing | language_nlp | 4,354 | 13 |
| UX-80 | Mood and intensity fit for "not now" filtering | ux_product | 21 | 0 |
| UX-81 | Read-mode recommendation | commerce_marketing | 3,400 | 12 |
| UX-82 | Comment quality re-ranking | ux_product | 9 | 0 |
| UX-83 | Live chat moderation triage | security_safety | 13,482 | 12 |
| UX-84 | Cold-start taste inference | commerce_marketing | 4 | 0 |
| UX-85 | Stopping-point detection for autoplay | ux_product | 9,492 | 4 |
| UX-86 | Perspective mix balancing | ux_product | 3,400 | 12 |
| UX-87 | Age-suitability scoring on shared devices | ux_product | 16 | 0 |
| UX-88 | Unlabeled control inference | ux_product | 15,453 | 18 |
| UX-89 | Reading-order repair | ux_product | 15,453 | 18 |
| UX-9 | Frustration detection from interaction stream | llm_rag_data | 12 | 0 |
| UX-90 | Verbosity control for screen readers | ux_product | 4,777 | 11 |
| UX-91 | Live-region announcement priority | geo_place | 4,777 | 11 |
| UX-92 | Essential-element selection for simplified mode | ux_product | 4,777 | 11 |
| UX-93 | Switch-access scan ordering | ux_product | 15,453 | 18 |
| UX-94 | Voice-control target disambiguation | speech_dialogue | 12,592 | 14 |
| UX-95 | Caption enrichment from ASR text | speech_dialogue | 9,486 | 4 |
| UX-96 | Sentence complexity gating for reading support | ux_product | 12 | 0 |
| UX-97 | Intended-target inference for imprecise touch | ux_product | 15,453 | 18 |
| UX-98 | Alt-text quality gate | ux_product | 4,777 | 11 |
| UX-99 | Photosensitivity risk pre-screen | ux_product | 4 | 0 |
| DATA-200 | Fine-grained entity mention typing (fan-out) | llm_rag_data | 9,105 | 12 |
| PHYS-202 | Chance-device outcome probability calibration | math_reasoning | 2,420 | 11 |
| DATA-203 | Stated-frequency and odds probability calibration | math_reasoning | 3,641 | 7 |
| DATA-201 | Domain-knowledge multiple-choice | knowledge_qa | 44,239 | 30 |
| BIZ-230 | Professional/legal/economic reasoning MCQ | finance_insurance | 17,159 | 4 |
| DATA-202 | Clinical knowledge MCQ | medicine_health | 7,923 | 12 |
| DEV-189 | Code weakness and security judgment | code_software | 9,006 | 11 |
| DEV-190 | Deterministic code-output prediction | code_software | 18,501 | 9 |
| DATA-204 | Arithmetic word-problem selection | math_reasoning | 40,025 | 11 |
| DATA-205 | Sarcasm and irony detection | language_nlp | 2,188 | 2 |
| DATA-206 | Sentiment classification | language_nlp | 6,035 | 1 |
| DATA-207 | Causal and counterfactual reasoning | math_reasoning | 25,199 | 7 |
| DATA-208 | Object tracking and belief attribution | math_reasoning | 22,125 | 21 |
| DATA-209 | Exact-set (all-and-only) selection | llm_rag_data | 1,187 | 1 |
| PHYS-203 | Chess move selection | games_puzzles | 17,790 | 17 |
| DATA-210 | Commonsense cloze and continuation | knowledge_qa | 28,596 | 21 |
| DATA-211 | Adversarial natural-language inference | security_safety | 41,182 | 15 |
| BIZ-232 | Contract provision classification (100-class) | legal_compliance | 499 | 2 |
| DATA-212 | Emotion classification | language_nlp | 26,518 | 31 |
| PHYS-201 | Dimensional-consistency checking | science_eng | 2,710 | 1 |
| DATA-213 | Constraint scheduling satisfiability | math_reasoning | 1,197 | 1 |
| DATA-214 | Time-series anomaly judgment | llm_rag_data | 1,381 | 1 |
| DATA-215 | Table-lookup predicate reasoning | llm_rag_data | 16,822 | 7 |
| DATA-216 | Paraphrase-invariance twin consistency | language_nlp | 11,026 | 14 |
| DATA-217 | Speech elicitation decision loops (letter/word/codebook/slot) | identity_meta | 954 | 3 |
| DATA-218 | Model identity anchoring under elicitation | identity_meta | 17 | 1 |
| DATA-219 | Real-world event probability forecasting (market-calibrated) | forecasting | 16,801 | 36 |
| DATA-220 | Literal-rule trap resistance | language_nlp | 11,990 | 1 |
| DATA-221 | Multi-hop rule chaining | llm_rag_data | 42,335 | 15 |
| DATA-222 | Probability estimation under uncertainty | math_reasoning | 10,059 | 5 |
| DATA-223 | Adversarial untrusted-text resistance | security_safety | 10,182 | 1 |
| DATA-224 | Temporal and window arithmetic | math_reasoning | 14,024 | 9 |
| DATA-225 | Passage relevance yes/no judgment (retrieval rerank) | llm_rag_data | 115,868 | 31 |
| DATA-226 | Best-passage selection over a candidate list (listwise rerank, incl. none) | llm_rag_data | 11,236 | 6 |
| DATA-227 | Pairwise passage preference (rerank) | llm_rag_data | 14,868 | 6 |
| DATA-228 | Graded passage relevance scoring | education | 62,748 | 21 |
| DATA-229 | Indirect-answer interpretation (yes/no from an indirect reply) | llm_rag_data | 19,589 | 0 |
| DATA-230 | Mathematical solution step verification | math_reasoning | 11,361 | 7 |
| DATA-231 | Multi-turn conversation safety rating | commerce_marketing | 8,519 | 12 |
| DATA-232 | Text-revision intent classification | language_nlp | 1,088 | 6 |
| DATA-233 | Protected-literal preservation in text edits | llm_rag_data | 483 | 1 |
| DATA-234 | Abductive explanation selection for a narrative | llm_rag_data | 488 | 1 |
| DATA-235 | Redaction edit review against a sanitisation policy | llm_rag_data | 5,090 | 0 |
| DATA-236 | Response claim support against a source (hallucination detection) | finance_insurance | 18,202 | 23 |
| DATA-237 | Adverse drug event mention detection | medicine_health | 3,639 | 11 |
| DATA-238 | Answerability and abstention: decide when the context does not support an answer | support_ops | 42,611 | 46 |
| DATA-239 | Multilingual natural-language inference | llm_rag_data | 6,521 | 13 |
| DATA-240 | Multilingual sentiment and review rating | commerce_marketing | 65,446 | 73 |
| DATA-241 | Multilingual topic classification | language_nlp | 15,008 | 22 |
| DATA-242 | Language identification | language_nlp | 2,566 | 4 |
| BIZ-233 | Phishing and scam message detection (email, SMS, chat) | security_safety | 11,400 | 11 |
| BIZ-234 | Resume to job-description fit | hr_workforce | 632 | 8 |
| BIZ-235 | Knowledge-article and macro selection for customer support | support_ops | 2,739 | 7 |
| DEV-191 | Infrastructure-as-code security rule compliance check | legal_compliance | 296 | 4 |
| DATA-243 | Few-shot and many-shot classification from labelled examples in the input | knowledge_qa | 3,528 | 10 |
| DATA-244 | Specification following: user-defined criteria, exceptions and counter-intuitive definitions | math_reasoning | 12,187 | 10 |
| DATA-245 | Very large label spaces via candidate shortlists (hierarchical or retrieved, with none-of-these) | llm_rag_data | 41,271 | 11 |
| DATA-246 | Multi-label set selection (select all that apply) | llm_rag_data | 101,944 | 29 |
| DATA-247 | Numeric and probabilistic estimation from real outcomes (calibrated bins) | math_reasoning | 18,027 | 31 |
| DATA-248 | Evidence attribution for a decision (supporting sentence or span) | math_reasoning | 76,649 | 28 |
| DATA-249 | Long-document questionnaire decisions (many questions per document) | llm_rag_data | 8,935 | 16 |
| DATA-250 | Counterfactual and contrast-set consistency (minimal edits that flip the answer) | math_reasoning | 24,421 | 21 |
| DEV-192 | Agent trajectory outcome and step verification from visible evidence | llm_rag_data | 18,038 | 3 |
| BIZ-236 | Control-framework crosswalk decisions (does a control satisfy a requirement) | business_ops | 1,130 | 5 |
| BIZ-237 | Patent examination outcome and classification | knowledge_qa | 11,353 | 12 |
| BIZ-238 | E-discovery responsiveness review | business_ops | 721 | 3 |
| DATA-251 | Tabular record outcome prediction (zero-shot and in-context, calibrated) | science_eng | 13,474 | 14 |
| DATA-252 | Human and institutional outcome prediction (persuasion, negotiation, review, community decisions) | llm_rag_data | 22,436 | 38 |
| DATA-253 | Personalised preference prediction from user history (recommendation and ranking) | hr_workforce | 53,548 | 71 |
| DATA-254 | Sequential action selection and plan validity | llm_rag_data | 22,861 | 7 |
| BIZ-239 | Healthcare and clinical decision support from public evidence | medicine_health | 2,937 | 9 |
| DATA-255 | Real-time game control from text-serialised state | games_puzzles | 9,011 | 21 |
| DATA-256 | Hidden-information and strategy game decisions (cards, poker, battles, drafts) | commerce_marketing | 38,703 | 20 |
| DATA-257 | Persona and audience response simulation | hr_workforce | 4,109 | 17 |
| DATA-258 | Engagement and virality prediction from pre-publication text | llm_rag_data | 3,093 | 3 |
| DATA-259 | Proactive conversation decisions: addressee, turn-taking, when to intervene, when to clarify, which reply | geo_place | 3,429 | 14 |
| DATA-260 | Goal-directed semantic navigation and association | llm_rag_data | 9,797 | 5 |
| DATA-261 | Document structure and style assignment | llm_rag_data | 3,046 | 2 |
| DATA-262 | Everyday activity choice from real behaviour logs | llm_rag_data | 2,870 | 4 |
| BIZ-240 | Crypto token and contract risk assessment | legal_compliance | 1,010 | 2 |
| DATA-263 | Multi-step temporal and numeric decisions inside records | math_reasoning | 29,129 | 18 |
| DATA-264 | Long rule documents with interacting conditions applied to a case | llm_rag_data | 16,288 | 20 |
| DATA-265 | Answer adequacy judging with subtle errors (math, code, constraints) | math_reasoning | 2,785 | 5 |
| DATA-266 | Multi-hop lookups across tables, notes and aliases | llm_rag_data | 7,422 | 4 |
| DATA-267 | AML typology identification from transaction networks | finance_insurance | 1,101 | 7 |
| DATA-268 | Crypto-asset address and transaction illicit-risk assessment | finance_insurance | 708 | 1 |
| DATA-269 | KYC beneficial ownership and control determination | finance_insurance | 677 | 1 |
| DATA-270 | PEP and relative-or-close-associate classification | finance_insurance | 285 | 1 |
| DATA-271 | AML regulatory reporting obligations (currency reports, SAR thresholds and deadlines, travel rule) | finance_insurance | 841 | 1 |
| DATA-272 | Sanctions and AML enforcement assessment (egregiousness, seriousness, penalty band, failures cited) | finance_insurance | 714 | 3 |
| DATA-273 | KYC customer risk rating and enhanced due diligence triggers | finance_insurance | 274 | 2 |
| DATA-274 | Payment fraud and money-mule detection | security_safety | 444 | 2 |
| BIZ-241 | AML SAR narrative completeness and accuracy review | finance_insurance | 195 | 2 |
| DATA-275 | Long-document decisions that combine distant facts | llm_rag_data | 10,724 | 8 |
| DATA-276 | Agent tool choice and call correctness | speech_dialogue | 56,021 | 26 |
| DATA-277 | Model routing and difficulty prediction | code_software | 846 | 2 |
| DATA-278 | Code change outcome prediction (merge, test failure) | code_software | 1,438 | 3 |
| DATA-279 | Vulnerability exploitation likelihood | security_safety | 1,117 | 2 |
| DATA-281 | Cognitive-bias resistance (framing, anchoring, sunk cost, decoy, base rates) | language_nlp | 814 | 4 |
| DATA-282 | Predicting human choices under risk | llm_rag_data | 8,940 | 6 |
| DATA-283 | Probabilistic updating and verbal probability | math_reasoning | 2,722 | 7 |
| DATA-284 | Causal versus correlational inference | math_reasoning | 6,288 | 4 |
| DATA-285 | Defeasible inference (does new information strengthen or weaken a conclusion) | commerce_marketing | 344 | 3 |
| DATA-286 | Higher-order theory of mind | llm_rag_data | 4,053 | 4 |
| DATA-287 | Spatial and temporal commonsense in text | knowledge_qa | 19,999 | 34 |
| DATA-288 | Dialogue and entity state tracking | speech_dialogue | 6,073 | 15 |
| DATA-289 | Moral and social-norm judgement with human vote distributions | llm_rag_data | 36,995 | 22 |
| DATA-290 | Argument quality and persuasiveness | llm_rag_data | 17,092 | 3 |
| DATA-291 | Humour and writing taste | llm_rag_data | 22,856 | 9 |
| DATA-292 | Consistency under user pushback (sycophancy resistance) | language_nlp | 14,056 | 20 |
| DATA-293 | Natural-language predicates over records (semantic filter, join, group, rank) | llm_rag_data | 4,764 | 1 |
| DATA-294 | Tutoring and teaching-move decisions | education | 14,386 | 18 |
| DATA-295 | Code and comment semantic consistency | data_quality | 1,782 | 5 |
| DATA-296 | Decisions on partial or streaming input (early intent, end of turn, wait or act) | speech_dialogue | 2,253 | 5 |
| DATA-297 | Individual-level response prediction (digital twins) | llm_rag_data | 7,856 | 12 |
| DATA-298 | Mapping language to perceptual choices (colour, emoji, music, visual style) | language_nlp | 2,880 | 7 |
| DATA-299 | Professional judgement in regulated work (triage acuity, eligibility, appeals, credit, environmental review) | medicine_health | 6,164 | 27 |
| DATA-300 | Many-field record completion (20 to 60 questions about one document) | llm_rag_data | 2,438 | 0 |
| DATA-301 | Evidence-bearing correction and robustness to inverted criteria and injected claims | finance_insurance | 7,667 | 3 |
| DATA-302 | Policy-conditioned safety decisions and severity | llm_rag_data | 23,263 | 18 |
| DATA-303 | Deep hierarchical classification by levels | llm_rag_data | 5,652 | 0 |
| DATA-304 | Spatial reasoning over serialised grids, boards and maps | llm_rag_data | 20,048 | 6 |
| DATA-305 | Everyday assistant decisions from real personal data | hr_workforce | 2,920 | 8 |
| DATA-306 | Incident and operations diagnosis | medicine_health | 1,864 | 1 |
| DATA-307 | Counting, magnitude and date-expression resolution | math_reasoning | 10,861 | 12 |
| JEVB-001 | Contested benchmark decisions, separately adjudicated diagnostics | Mapping pending | 0 | 0 |
| DEV-193 | Text-environment state after an action log (holding, location, open containers) | geo_place | 41,008 | 35 |
| DATA-308 | Aspect-category sentiment fan-out over a review (all categories, not discussed) | language_nlp | 33,643 | 16 |
| DATA-329 | Multi-turn escalation toward refused content | support_ops | 22,429 | 27 |
| DATA-309 | Stance toward a stated target (favour, against, neutral) | language_nlp | 30,634 | 22 |
| DATA-310 | Option-order invariance: the same question in several option orders | llm_rag_data | 13,115 | 42 |
| DEV-194 | Web element to act on among 10-254 indexed page elements | llm_rag_data | 8,455 | 2 |
| DATA-311 | Choice and per-option yes/no coherence on one state | llm_rag_data | 16,941 | 18 |
| DATA-312 | Formal deduction over stated facts and rules (true, false, unknown) | commerce_marketing | 20,898 | 11 |
| BIZ-242 | Court outcome prediction from case facts (ECtHR, Swiss FSC, ILDC) | legal_compliance | 10,471 | 12 |
| DATA-313 | Cross-lingual and multilingual paraphrase identification | language_nlp | 17,307 | 13 |
| DATA-314 | Writing defects in machine-written prose | llm_rag_data | 3,347 | 4 |
| DATA-315 | Quantifiers over sibling judgements or record lists (all, some, none, at least k) | science_eng | 14,250 | 6 |
| DATA-316 | Group and member preference over candidate consensus statements | llm_rag_data | 11,632 | 9 |
| DATA-317 | Chess position value for the side to move (engine win-probability bands) | math_reasoning | 8,277 | 8 |
| DATA-394 | One yes/no per supplied rubric criterion or checklist requirement on one response (HealthBench, ProfBench, InF | medicine_health | 9,624 | 15 |
| DATA-395 | Classify a software requirement (SRS line or GitHub issue) by quality attribute (ISO/IEC 25010-style), functio | math_reasoning | 13,018 | 11 |
| DATA-319 | Clinical calculator score from values in a patient note | medicine_health | 12,725 | 11 |
| DATA-320 | Resolved real-world forecasting questions | forecasting | 7,570 | 24 |
| BIZ-243 | Convicted charge from Chinese criminal case facts | business_ops | 12,612 | 4 |
| DEV-208 | Alert / anomaly triage over long system-log excerpts (which entries, windows, blocks were flagged by the datas | infra_ops | 12,468 | 19 |
| BIZ-244 | Market reaction grade to a news flash | education | 11,882 | 16 |
| DATA-321 | Chord of a span in real music from melody and notes | llm_rag_data | 11,542 | 6 |
| DATA-396 | Medical exam items (USMLE, AIIMS/NEET PG) | medicine_health | 11,525 | 7 |
| DEV-195 | Context compaction: keep or drop each tool-output block of an agent history | llm_rag_data | 0 | 0 |
| DATA-322 | Two-party chat reading: the other speaker’s emotion and act, the reader’s own view | speech_dialogue | 11,119 | 12 |
| DATA-323 | Slot presence, value and count in voice-assistant utterances | math_reasoning | 6,833 | 11 |
| BIZ-245 | Sentiment toward a named company in financial text | finance_insurance | 10,405 | 8 |
| DATA-324 | Count of options that apply (select-all-that-apply consistency) | math_reasoning | 10,774 | 2 |
| DATA-325 | Time-series horizon questions with empirical base-rate targets | llm_rag_data | 0 | 0 |
| BIZ-251 | Classify a real procurement record (CPV / ДК021 division, contract nature, declared category, contract type, s | finance_insurance | 10,230 | 10 |
| DATA-326 | Packed records: 20-80 labelled records in one JSON state, one question per record | code_software | 9,384 | 0 |
| DATA-327 | Sentiment toward an aspect term in a review | language_nlp | 17,146 | 24 |
| DATA-328 | Single-message jailbreak intent under a deployment card | security_safety | 9,926 | 4 |
| DATA-330 | Live secret vs placeholder, hash or public value in context | security_safety | 191 | 1 |
| DEV-196 | First code line that breaks a stated rule | code_software | 9,155 | 8 |
| DATA-332 | Exam reading comprehension with all questions on one passage | llm_rag_data | 7,113 | 0 |
| DATA-331 | Negation-sensitive relevance (twin queries over the same list) | llm_rag_data | 1,651 | 0 |
| BIZ-252 | Court / chamber / case type / decision type routing from the reasoning text | legal_compliance | 8,911 | 8 |
| DATA-333 | Exact chord naming from a voiced bar | speech_dialogue | 8,889 | 9 |
| DEV-197 | Which real review comment was written on a code hunk | code_software | 8,642 | 0 |
| BIZ-253 | Hawkish/dovish stance, forward-looking and certainty of central bank sentences (WCB, FOMC) | finance_insurance | 8,270 | 6 |
| BIZ-254 | US-GAAP element for a numeral in a filing sentence (FiNER-139) | business_ops | 7,993 | 8 |
| BIZ-255 | Reading long 10-K statement / MD&A sections | business_ops | 7,921 | 18 |
| DATA-397 | TREC CT judgement (eligible / excluded / not relevant with the track definitions), eligible and target-conditi | medicine_health | 7,832 | 0 |
| DEV-209 | Reviewing a long coding-agent run record | llm_rag_data | 7,807 | 10 |
| DEV-210 | Coding-agent permission decision for a full shell script (allow, ask, deny) | llm_rag_data | 7,683 | 8 |
| DATA-351 | Phishing link and header evidence over persuasive prose | security_safety | 7,381 | 8 |
| DATA-335 | Discourse relation between two text units | education | 6,363 | 9 |
| DATA-398 | Differential diagnosis from a patient symptom and history record (DDXPlus) | medicine_health | 7,376 | 8 |
| BIZ-246 | Editorial section of a dated news item | ux_product | 7,282 | 2 |
| DATA-399 | Classify a chemical reaction (reactants to product) into a named reaction class, or pick the reported product | science_eng | 7,216 | 4 |
| DATA-336 | Native-speaker naturalness and spelling ratings of an utterance | commerce_marketing | 2 | 0 |
| DATA-318 | A yes/no question and its negation on one state | llm_rag_data | 6,194 | 5 |
| BIZ-256 | Trading and market-state decisions (real public sources) | finance_insurance | 6,988 | 10 |
| DATA-337 | Over-defence: benign text that contains instructions for a person | llm_rag_data | 6,880 | 4 |
| DATA-338 | Missing hop: which step or page is still needed | llm_rag_data | 6,150 | 0 |
| DEV-211 | Coding-agent prompt-cache reuse vs rebuild (token cost break-even over remaining turns) | code_software | 6,702 | 5 |
| DEV-212 | Which tool-output chunks a coding agent must keep visible for the next step | llm_rag_data | 6,635 | 0 |
| BIZ-257 | Court decision outcome from the reasoning with the operative part removed | legal_compliance | 6,530 | 8 |
| DATA-400 | Declarative statement about the record, Yes/No | finance_insurance | 6,017 | 0 |
| DEV-198 | Does code break a plain-English rule | code_software | 6,377 | 3 |
| BIZ-258 | Reading long court opinions | legal_compliance | 6,345 | 12 |
| DATA-339 | Follow a stated relation exactly k steps | llm_rag_data | 6,166 | 4 |
| DEV-213 | Cost-priced model routing for a coding-agent step (cheapest adequate tier) | finance_insurance | 6,155 | 5 |
| DATA-401 | Supported / Contradicted / Not stated for one specific claim | finance_insurance | 5,614 | 0 |
| DATA-402 | Alignment-failure monitor: privacy leak in an AI transcript | infra_ops | 6,028 | 20 |
| DATA-403 | Native-speaker grammatical acceptability or error-free judgement of a sentence (CoLA-style sets from linguisti | speech_dialogue | 6,034 | 12 |
| DATA-340 | Concept commonsense questions (CommonsenseQA) | knowledge_qa | 4,059 | 3 |
| DATA-341 | Social commonsense: motivations, reactions and next actions | knowledge_qa | 3,793 | 3 |
| DATA-404 | Compliance of a record with a short stated policy (2-5 rules) | legal_compliance | 5,376 | 0 |
| DATA-342 | Does any candidate in a list answer the query | llm_rag_data | 5,021 | 3 |
| DEV-199 | Did a reviewer comment on this code hunk | code_software | 0 | 0 |
| DATA-343 | Ordinal scale listed in reverse or with level texts | commerce_marketing | 3,519 | 25 |
| DATA-344 | Withheld evidence: deciding field removed, empirical target | medicine_health | 5,712 | 0 |
| DATA-346 | Commonsense inference over a passage, several questions per passage | llm_rag_data | 3,811 | 0 |
| DATA-345 | Physical commonsense: which method achieves a goal | science_eng | 3,760 | 1 |
| DATA-405 | Level of a learning resource (reading text or sentence) on the CEFR or a course-level scale, for routing mater | education | 2,869 | 5 |
| BIZ-259 | Climate relevance, TCFD element, risk vs opportunity, specificity and commitments in annual-report paragraphs | science_eng | 5,593 | 7 |
| DATA-406 | Alignment-failure monitor: power-seeking behaviour in an AI transcript | infra_ops | 5,526 | 12 |
| BIZ-260 | Classify a corporate current report (8-K) by the SEC items it reports, from the body text with headings remove | business_ops | 5,515 | 3 |
| DATA-407 | Code-verifiable instruction constraints (word/sentence counts, formats, keywords) checked on a real response | math_reasoning | 5,491 | 9 |
| DATA-408 | Next action under a stated rule | llm_rag_data | 5,035 | 0 |
| DEV-214 | Browser action selection (DOM) (synthetic app scenarios) | code_software | 5,013 | 0 |
| DATA-409 | Urgency or priority tier with defined tiers | llm_rag_data | 4,785 | 0 |
| DATA-347 | Next chord in a real song chart | llm_rag_data | 5,200 | 0 |
| DATA-348 | Reading with commonsense inference (QuAIL, LSAT-RC) | knowledge_qa | 5,086 | 11 |
| DEV-215 | Which repository files a change will touch (observed from the merged patch) | code_software | 5,084 | 9 |
| DATA-410 | Send the record to one of several described queues | llm_rag_data | 4,631 | 0 |
| BIZ-261 | Procurement deadline arithmetic and open/closed checks against an as-of date | finance_insurance | 4,993 | 5 |
| DEV-216 | Shared-work coordination: read/write typing and collisions between parallel agent steps | llm_rag_data | 4,944 | 8 |
| DEV-217 | Alignment-failure monitor: prompt injection followed by an AI agent | security_safety | 4,904 | 8 |
| DATA-411 | Code an injury narrative to OIICS event, nature, part of body and source | llm_rag_data | 4,886 | 8 |
| DEV-218 | Final status of a support thread, asker confirmation, rejection as invalid, solving reply | commerce_marketing | 4,876 | 8 |
| DATA-349 | Kinship composition over 2-4 hops | llm_rag_data | 3,274 | 4 |
| DATA-350 | Crowd NLI vote distribution (100 raters) | language_nlp | 664 | 3 |
| BIZ-262 | Whole-contract review | legal_compliance | 4,754 | 12 |
| DATA-412 | Order the steps of a laboratory protocol (first, last, next, before/after) | llm_rag_data | 4,750 | 0 |
| DEV-219 | Label a new GitHub issue as bug, feature or question | code_software | 4,488 | 3 |
| DATA-413 | Value of a record attribute from described values | math_reasoning | 4,114 | 0 |
| DATA-414 | Ordinal quality/completeness/clarity score (type score) | commerce_marketing | 4,086 | 0 |
| DEV-200 | Agent trajectory hijack and the step that carried the instruction | llm_rag_data | 3,942 | 0 |
| DATA-415 | Ordinal risk score, levels defined in the instructions (type score) | llm_rag_data | 4,012 | 0 |
| BIZ-263 | Privacy-policy segment data-practice tagging (OPP-115 scheme) with annotator vote counts | math_reasoning | 635 | 0 |
| DATA-416 | Moderation, spam, phishing and scam detection (real public sources) | security_safety | 4,167 | 9 |
| BIZ-264 | Law type, subject category, era-year conversion, article count | math_reasoning | 4,064 | 7 |
| BIZ-265 | Apply a procedure rule or legal basis to a record (open vs restricted access, statutory thresholds, legal basi | legal_compliance | 4,021 | 4 |
| DATA-417 | The same response scored under several different supplied rubrics (UltraFeedback aspects) | code_software | 0 | 0 |
| DATA-418 | Quantitative NLI (EQUATE-style, NumGLUE type 7) with the three relation definitions in criteria | science_eng | 4,000 | 0 |
| BIZ-266 | Argumentative role (conclusion, definition, subsumption, other) of a numbered sentence in a German court decis | legal_compliance | 3,897 | 12 |
| BIZ-267 | Lending Club grade, rate band and charge-off from the application | education | 3,894 | 6 |
| DATA-419 | Alignment-failure monitor: overconfident or miscalibrated AI answer | science_eng | 3,864 | 14 |
| BIZ-268 | Trading and market-state decisions (synthetic app scenarios) | finance_insurance | 3,868 | 0 |
| DATA-352 | Graded semantic similarity score | education | 3,814 | 9 |
| DATA-420 | Aspect-level sentiment fan-out on one review (ParsiNLU food/movie aspects, IndoNLU CASA and HoASA, UIT-ViSFD) | language_nlp | 3,796 | 4 |
| DATA-353 | How many shown documents an answer depends on | llm_rag_data | 2,754 | 1 |
| DATA-421 | DDI 2013 type per mention pair (mechanism, effect, advice, int, none) and the count of interacting pairs | math_reasoning | 3,083 | 3 |
| DATA-354 | Defective options: typo in gold or gold missing | language_nlp | 0 | 0 |
| DATA-355 | Figurative language: metaphor, idiom and simile meaning | llm_rag_data | 3,611 | 1 |
| BIZ-269 | Which Form 10-K item an excerpt comes from | business_ops | 3,597 | 6 |
| BIZ-270 | Insurance regulator complaint coding and outcome | finance_insurance | 3,596 | 4 |
| DATA-422 | Twin state with the same documents re-ordered (decisive text moved to a different position) | code_software | 2,323 | 3 |
| DATA-423 | Clinical note specialty filing (MTSamples) | medicine_health | 3,542 | 4 |
| DATA-424 | Alignment-failure monitor: deception in an AI transcript | security_safety | 3,514 | 16 |
| DATA-356 | Dialogue act of a printed message | speech_dialogue | 1,718 | 6 |
| DATA-357 | Symbolic procedures (Dyck, navigation, counting, arithmetic, sorting) | math_reasoning | 3,499 | 1 |
| BIZ-271 | Document routing, extraction and field-level data QA (real public sources) | knowledge_qa | 3,487 | 8 |
| DEV-220 | Alignment-failure monitor: reward hacking or specification gaming by an AI agent | infra_ops | 3,416 | 10 |
| DEV-221 | File sensitivity and trust routing for a coding agent (secrets, credentials, config) | security_safety | 3,401 | 4 |
| DATA-358 | School science questions | education | 3,395 | 3 |
| DATA-359 | Candidate stance on a Swiss ballot proposal (de, fr) | gov_civic | 3,357 | 10 |
| DATA-360 | Pass a stated probability or base rate through a decision | math_reasoning | 3,345 | 9 |
| DATA-361 | Sentiment toward a target phrase in Japanese financial text | finance_insurance | 3,350 | 0 |
| DEV-222 | Did the ticket or incident meet its time target | support_ops | 3,291 | 7 |
| DEV-223 | General agent decision layer (mixed route/gate/check in one harness) (synthetic app scenarios) | travel_hospitality | 2,844 | 0 |
| DATA-425 | Ranking, curation and news/paper radar (real public sources) | llm_rag_data | 3,222 | 7 |
| DEV-201 | LLM call trace behaviours over call metadata (latency, tokens, cost, sampling) | infra_ops | 3,220 | 5 |
| BIZ-247 | Observed outcome of a closed parliamentary petition | gov_civic | 2,599 | 6 |
| DATA-362 | Argument reasoning: warrants, fallacies, verifiability | llm_rag_data | 3,185 | 1 |
| DEV-224 | Tool, skill and MCP selection (synthetic app scenarios) | code_software | 2,405 | 0 |
| DATA-426 | Turn-based, board, card and strategy game moves (real public sources) | speech_dialogue | 3,122 | 10 |
| DATA-363 | Pragmatics: implied meaning, formality, valence | commerce_marketing | 3,121 | 3 |
| DATA-427 | Chat reading: intent, emotion, relationship and reply ranking (synthetic app scenarios) | speech_dialogue | 2,858 | 0 |
| DATA-428 | Judge a computed materials property band or threshold from a structure description | travel_hospitality | 3,083 | 6 |
| DATA-364 | Variable value after line n of a program | code_software | 3,060 | 3 |
| BIZ-272 | Gold headline dimensions for a trading desk | finance_insurance | 2,992 | 6 |
| DATA-365 | Stock level after transaction n of a ledger | finance_insurance | 2,981 | 3 |
| DEV-225 | Which vendor components an incident affected, from the first public notice | finance_insurance | 2,960 | 6 |
| DATA-429 | Route an AI transcript to its alignment-failure type | speech_dialogue | 2,956 | 10 |
| DATA-366 | How many listed passages help a reasoning-intensive question | llm_rag_data | 1,898 | 1 |
| BIZ-273 | Headnote-to-decision match and cited-provision identification | business_ops | 2,835 | 5 |
| BIZ-274 | Incident, SRE, SOC and pentest-finding triage (real public sources) | infra_ops | 2,786 | 7 |
| DATA-367 | Claim truth (fact checking) | finance_insurance | 2,758 | 3 |
| DATA-368 | Emotional-support strategy of the next supporter turn | speech_dialogue | 2,726 | 3 |
| DEV-226 | Memory, retrieval and passage relevance (RAG rerank) (real public sources) | llm_rag_data | 2,704 | 0 |
| DEV-227 | Model, effort and subagent routing (synthetic app scenarios) | llm_rag_data | 2,285 | 0 |
| DATA-430 | Feed filtering: AI-slop, ads, bait and topic folding (synthetic app scenarios) | language_nlp | 2,286 | 0 |
| DEV-228 | Desktop and mobile computer use (AX tree, OCR) (synthetic app scenarios) | language_nlp | 2,381 | 0 |
| DATA-369 | Count of candidates meeting a relevance rule | math_reasoning | 1,980 | 1 |
| DEV-202 | Python API availability by version | code_software | 2,549 | 0 |
| DATA-431 | Which required checklist item is missing, or complete | llm_rag_data | 2,347 | 0 |
| DATA-370 | Count band of relevant candidates in a list | math_reasoning | 3 | 0 |
| DATA-432 | Read or compute a quantity and choose its band | science_eng | 2,311 | 0 |
| DATA-433 | Which programme, tier or offer the record qualifies for, or none | code_software | 2,301 | 0 |
| DEV-229 | Which subdirectory instruction files (AGENTS.md, CLAUDE.md) apply to an edited file | llm_rag_data | 2,458 | 11 |
| DEV-230 | Evaluate or pick a spreadsheet formula (SUM/AVERAGE/MAX/MIN/COUNTIF/SUMIF) over a real table shown as a sheet | math_reasoning | 2,355 | 6 |
| DEV-231 | Apply a W3C ACT accessibility rule to HTML/CSS/SVG: passed / failed / inapplicable | ux_product | 2,255 | 2 |
| DATA-434 | Which entity in the record meets a stated condition | llm_rag_data | 2,112 | 0 |
| DATA-435 | National identifier checksum validity (country-specific ID rules) | math_reasoning | 2,204 | 4 |
| DEV-232 | Planned vs actual implementation windows of real change records (temporal comparison) | finance_insurance | 2,133 | 4 |
| DATA-371 | Stance of a reply or headline toward a rumour or claim | finance_insurance | 2,122 | 2 |
| DATA-372 | Dialogue act of an utterance (Switchboard, DailyDialog) | speech_dialogue | 1,061 | 1 |
| DATA-373 | How-to goal and step order | llm_rag_data | 2,122 | 0 |
| DATA-436 | Order-of-magnitude and unit sense from real country statistics | math_reasoning | 2,057 | 10 |
| DEV-233 | Context compaction and tool-output pruning (synthetic app scenarios) | code_software | 1,678 | 0 |
| DATA-437 | Opinion and sentiment aggregation over comments and reviews (real public sources) | language_nlp | 2,002 | 10 |
| DEV-234 | Resolver closure code and cause traced to another configuration item | finance_insurance | 1,996 | 4 |
| DATA-438 | Whether two parts of the record agree, or the kind of mismatch | knowledge_qa | 1,821 | 0 |
| DATA-439 | Real-time game and vehicle control from text state (synthetic app scenarios) | science_eng | 1,864 | 0 |
| BIZ-275 | Payment terms (share, days, day type, advance vs after performance) and contract duration | finance_insurance | 1,953 | 3 |
| DEV-235 | Which merged PR resolved a problem report | code_software | 1,932 | 3 |
| DEV-236 | Model, effort and subagent routing (real public sources) | llm_rag_data | 1,905 | 4 |
| DATA-440 | Applicable VAT rate for a transaction under a country rule | finance_insurance | 1,892 | 4 |
| DEV-203 | Data leaving its boundary in an agent trajectory | llm_rag_data | 1,825 | 12 |
| DATA-374 | No evidence about a chance outcome: uniform probabilities | math_reasoning | 1,882 | 4 |
| DATA-441 | CPT section / surgical subsection and ICD-10-CM chapter from the note | llm_rag_data | 1,858 | 2 |
| DATA-442 | Registered age and sex limits (computed twice) and per-criterion met / not met / unknown (two-model agreement) | medicine_health | 1,854 | 0 |
| DEV-237 | Semantic lint, code review and commit checks (synthetic app scenarios) | code_software | 800 | 0 |
| DEV-238 | Memory, retrieval and passage relevance (RAG rerank) (synthetic app scenarios) | llm_rag_data | 1,729 | 0 |
| DEV-239 | What context a coding agent must pass to a sub-agent | llm_rag_data | 1,840 | 0 |
| DEV-240 | Whether a proposed agent subgoal duplicates work already done or in flight | llm_rag_data | 1,831 | 0 |
| DEV-241 | Agent-trace and LLM-output evaluation, output guardrails (synthetic app scenarios) | security_safety | 1,653 | 0 |
| DATA-375 | Opinion span for an aspect in a review | llm_rag_data | 1,790 | 0 |
| DATA-376 | Where the evidence sits (table, notes or both) | math_reasoning | 1,633 | 0 |
| DATA-377 | Implicit multi-step yes/no questions | knowledge_qa | 1,697 | 1 |
| DATA-378 | Frequency judgements with rater spread | llm_rag_data | 0 | 0 |
| DATA-443 | Locale grammar and calendar rules (plural forms, first day of week) | language_nlp | 1,684 | 4 |
| DEV-242 | How an issue was eventually closed (completed / not planned / duplicate) | code_software | 1,676 | 3 |
| DATA-444 | Which described reply template to send | llm_rag_data | 1,546 | 0 |
| DATA-445 | Personal organisation (tabs, references, files) (real public sources) | hr_workforce | 1,619 | 7 |
| DATA-446 | Native Chinese exam reading comprehension (C3) | knowledge_qa | 1,608 | 8 |
| BIZ-276 | Support, inbox, email and notification triage (real public sources) | ux_product | 1,594 | 5 |
| DATA-447 | Order, date relation or elapsed-time band | llm_rag_data | 1,462 | 0 |
| DATA-448 | Graded semantic similarity (JSTS Japanese, MahaSTS Marathi) | education | 1,506 | 6 |
| BIZ-277 | Numeric claim vs reported figure in analyst sentences | finance_insurance | 1,499 | 1 |
| DATA-449 | Multilingual politeness of short editor-discussion excerpts | ux_product | 1,489 | 5 |
| DATA-450 | Feed filtering: AI-slop, ads, bait and topic folding (real public sources) | language_nlp | 1,433 | 0 |
| DATA-451 | Same or different sense of a word in two native sentences (WiC-ITA, RUSSE) | code_software | 1,395 | 3 |
| DEV-243 | Agent action safety and approval gates (synthetic app scenarios) | llm_rag_data | 1,034 | 0 |
| DATA-452 | Writing-quality scoring and prose linting (synthetic app scenarios) | code_software | 1,198 | 0 |
| DEV-244 | Conditional instruction loading: does a glob-scoped rule file apply to this edit | code_software | 1,333 | 0 |
| DATA-380 | States listed in a recall distribution pattern | medicine_health | 1,319 | 6 |
| DATA-453 | Face act of a dialogue turn (Brown and Levinson face framework as adapted by Dutt et al | speech_dialogue | 1,275 | 4 |
| DATA-454 | Real-time game and vehicle control from text state (real public sources) | science_eng | 1,314 | 3 |
| DATA-455 | Planning, graph navigation and constraint checking (real public sources) | legal_compliance | 1,289 | 4 |
| DATA-381 | Expected answer type of a search query | llm_rag_data | 1,247 | 1 |
| BIZ-248 | Case-law holding selection | business_ops | 1,219 | 3 |
| BIZ-278 | Match a GDPR case (facts) to the holding actually given, among holdings of other decisions | legal_compliance | 1,210 | 2 |
| DEV-245 | Agent progress, stuck-loop and done-claim checks (synthetic app scenarios) | finance_insurance | 1,153 | 0 |
| DATA-456 | DDXPlus severity (1 most severe .. 5 least severe) of the patient's condition | medicine_health | 1,175 | 1 |
| DATA-457 | Award-criterion type, price-only check, quality weight arithmetic | finance_insurance | 1,174 | 2 |
| BIZ-279 | Sales, marketing, SEO, recruiting and startup-idea scoring (synthetic app scenarios) | finance_insurance | 1,120 | 0 |
| DATA-458 | Ranking, curation and news/paper radar (synthetic app scenarios) | llm_rag_data | 1,104 | 0 |
| BIZ-280 | Support, inbox, email and notification triage (synthetic app scenarios) | ux_product | 1,030 | 0 |
| DATA-382 | Which target a stance is expressed toward | language_nlp | 1,122 | 1 |
| DATA-459 | Qualitative consequence of two stated quantities (NumGLUE type 3, QuaRel-style) | science_eng | 1,116 | 0 |
| DEV-204 | Vulnerability label for a C function | security_safety | 1,061 | 1 |
| DATA-383 | Next-turn coherence in dialogue | speech_dialogue | 1,057 | 1 |
| DATA-384 | Offensive-language detection outside English | code_software | 1,059 | 3 |
| DATA-385 | Expected answer type of a question | llm_rag_data | 1,062 | 0 |
| DATA-386 | Story completion in nine languages | knowledge_qa | 1,061 | 1 |
| DATA-387 | Why a paper is cited | language_nlp | 1,058 | 0 |
| DATA-460 | Robotics, drones and physical control (synthetic app scenarios) | science_eng | 1,009 | 0 |
| DEV-246 | Semantic lint, code review and commit checks (real public sources) | code_software | 1,043 | 3 |
| DATA-334 | Comparison and ordering of quantities and dates | science_eng | 926 | 3 |
| DATA-461 | MTSamples title among same-specialty titles | llm_rag_data | 1,014 | 1 |
| BIZ-281 | Incident, SRE, SOC and pentest-finding triage (synthetic app scenarios) | infra_ops | 987 | 0 |
| DEV-247 | Intent from short input: UI morphing, CLI typo, IME and history ranking (synthetic app scenarios) | language_nlp | 856 | 0 |
| DEV-248 | Semantic SQL, dataframes and warehouse predicates (synthetic app scenarios) | code_software | 915 | 0 |
| DEV-249 | Better next step at a real agent decision fork, judged by what the recorded runs showed | llm_rag_data | 992 | 12 |
| BIZ-282 | Sales, marketing, SEO, recruiting and startup-idea scoring (real public sources) | finance_insurance | 977 | 3 |
| DATA-379 | Same text, different target: target-relative judgement | code_software | 835 | 1 |
| DATA-462 | Smart home, voice pipelines and device control (synthetic app scenarios) | infra_ops | 624 | 0 |
| DATA-388 | How many listed fields a message gives a value for | code_software | 397 | 1 |
| DEV-250 | Semantic grep and code/file search by meaning (synthetic app scenarios) | llm_rag_data | 830 | 0 |
| DATA-463 | Elapsed time between filing and decision | llm_rag_data | 881 | 1 |
| DATA-389 | Word-pair analogies | llm_rag_data | 867 | 1 |
| DEV-251 | Name the WCAG 2.2 success criterion a failing example puts at risk | knowledge_qa | 792 | 3 |
| DEV-205 | Which listed call returns a given value | commerce_marketing | 795 | 1 |
| DATA-464 | Planning, graph navigation and constraint checking (synthetic app scenarios) | llm_rag_data | 788 | 0 |
| DATA-465 | Native paraphrase pairs (MahaParaphrase Marathi, ParsiNLU natural query pairs Persian) | language_nlp | 754 | 2 |
| DEV-252 | Dev-loop micro-decisions (test selection, debugger step, compiler inlining, CI health, PR routing) (synthetic app scenarios) | medicine_health | 614 | 0 |
| DATA-466 | Entity relevance | llm_rag_data | 743 | 1 |
| DATA-467 | German Sustainability Code criterion addressed by a sentence of a sustainability report | geo_place | 731 | 4 |
| DATA-468 | Generation by selection (music, pixels, UI, video, GIFs, tokens, numbers) (synthetic app scenarios) | ux_product | 674 | 0 |
| DATA-469 | Whether a Polish consumer-contract clause is abusive (UOKiK register) | legal_compliance | 694 | 1 |
| DEV-253 | Intent from short input: UI morphing, CLI typo, IME and history ranking (real public sources) | language_nlp | 690 | 0 |
| DATA-470 | Party of a Norwegian parliament speaker from the speech text | gov_civic | 685 | 3 |
| DEV-254 | Which instructions must stay pinned in a coding agent context across turns | travel_hospitality | 675 | 5 |
| DATA-471 | Procurement selection-criterion category (EU Directive 2014/24) | finance_insurance | 665 | 1 |
| DATA-472 | UIT-ViCTSD: is a Vietnamese news comment constructive | llm_rag_data | 664 | 0 |
| DEV-206 | Is a customer request blocked by policy | llm_rag_data | 295 | 0 |
| DATA-473 | Five-year period in which an Italian historical document was written (ordered scale) | llm_rag_data | 645 | 3 |
| BIZ-283 | Finance back office (transactions, receipts, tax forms, claims, AML) (real public sources) | finance_insurance | 635 | 0 |
| BIZ-284 | Health, education, legal and compliance domain decisions (real public sources) | medicine_health | 624 | 4 |
| DEV-255 | Security: injection filters, malicious code, bot detection, tool authorization (real public sources) | llm_rag_data | 622 | 0 |
| DATA-474 | Whether text B summarises text A (Polish Summaries Corpus) | language_nlp | 599 | 1 |
| BIZ-285 | Finance back office (transactions, receipts, tax forms, claims, AML) (synthetic app scenarios) | finance_insurance | 476 | 0 |
| DEV-256 | Was an infra issue closed as completed or not planned | code_software | 593 | 1 |
| DEV-257 | Was a change later linked to incidents (change risk outcome, natural rare positive) | infra_ops | 585 | 1 |
| BIZ-286 | Document routing, extraction and field-level data QA (synthetic app scenarios) | knowledge_qa | 425 | 0 |
| DATA-475 | Which test the harness recorded as fail-to-pass for a fix | code_software | 576 | 2 |
| DEV-258 | Security: injection filters, malicious code, bot detection, tool authorization (synthetic app scenarios) | llm_rag_data | 475 | 0 |
| DATA-390 | Does an utterance contain English words (code-switching) | speech_dialogue | 2 | 0 |
| BIZ-287 | Health, education, legal and compliance domain decisions (synthetic app scenarios) | medicine_health | 482 | 0 |
| DATA-476 | Twin state with the decisive document/summary/episode removed | language_nlp | 332 | 3 |
| DATA-477 | Turn-based, board, card and strategy game moves (synthetic app scenarios) | speech_dialogue | 369 | 0 |
| DATA-478 | Moderation, spam, phishing and scam detection (synthetic app scenarios) | security_safety | 501 | 0 |
| DATA-479 | Smart home, voice pipelines and device control (real public sources) | infra_ops | 488 | 3 |
| DATA-480 | Decide which retrieved tracker issue (if any) a user review should be linked to | code_software | 482 | 1 |
| BIZ-249 | Product-details field value stated in a listing | travel_hospitality | 464 | 7 |
| DATA-481 | Robotics, drones and physical control (real public sources) | science_eng | 414 | 2 |
| DATA-482 | Opinion and sentiment aggregation over comments and reviews (synthetic app scenarios) | language_nlp | 382 | 0 |
| DATA-483 | Personal organisation (tabs, references, files) (synthetic app scenarios) | hr_workforce | 403 | 0 |
| DATA-484 | Research literature screening and review extraction (real public sources) | llm_rag_data | 388 | 3 |
| DATA-485 | Priority or severity level maintainers gave an issue (ordered levels) | code_software | 367 | 2 |
| BIZ-250 | Category path name of a product listing | travel_hospitality | 354 | 2 |
| DATA-486 | NPC, agent-society and narrative choice (synthetic app scenarios) | science_eng | 291 | 0 |
| DATA-487 | Research literature screening and review extraction (synthetic app scenarios) | llm_rag_data | 288 | 0 |
| DATA-488 | CLUE CSL: are all listed keywords the paper's genuine keywords (source pairs each abstract with a genuine and | llm_rag_data | 242 | 1 |
| DATA-489 | Identify the culprit(s) of a whole mystery story, including name-swapped copies | llm_rag_data | 233 | 6 |
| DATA-391 | Base rate of a value among other records when a record is redacted | infra_ops | 225 | 0 |
| DEV-259 | Browser action selection (DOM) (real public sources) | code_software | 205 | 4 |
| DATA-392 | Which information a message asks for | commerce_marketing | 123 | 0 |
| DATA-490 | Word and party games that judge a human (synthetic app scenarios) | games_puzzles | 136 | 0 |
| DEV-260 | Desktop and mobile computer use (AX tree, OCR) (real public sources) | language_nlp | 105 | 0 |
| DEV-261 | Agent progress, stuck-loop and done-claim checks (real public sources) | finance_insurance | 91 | 1 |
| DEV-207 | Execution under an older Python version | code_software | 58 | 2 |
| DEV-262 | Dev-loop micro-decisions (test selection, debugger step, compiler inlining, CI health, PR routing) (real public sources) | medicine_health | 53 | 5 |
| DATA-393 | Does an abstract support or contradict a scientific claim | finance_insurance | 47 | 0 |