Our Methodology
Jump to: Data reuse · Revenue readiness
Data Source
All data comes from the United States Sentencing Commission (USSC), an independent agency of the judicial branch that sets federal sentencing policy and collects data on every federal offender sentenced in U.S. district courts. The USSC publishes annual public-use datafiles covering each fiscal year (October 1 – September 30).
Our database aggregates FY2015 through FY2025 - 728,223 federal sentencing
records across 93 judicial districts in this extract and 30 offense types. District names follow the USSC DISTRICT codebook (not the CIRCDIST Sourcebook-order recode).
Cross-references for federal sentencing context include the Administrative Office of the U.S. Courts statistical reports (caseload and judicial workload statistics by district), the Bureau of Justice Statistics (BJS national-scope criminal justice analyses), and the Federal Judicial Center Integrated Database (case-level docket data). For the binding sentencing rules themselves, the USSC Guidelines Manual is authoritative.
Data Format and Extraction
USSC annual datafiles are distributed in fixed-width SPSS format (FY2015 through FY2023) and CSV format (FY2024–FY2025). Each annual file contains tens of thousands of individual sentencing records. We extract the following variables for each case record:
- Federal judicial district
- Primary offense guideline category
- Sentence length in months
- Guideline minimum and maximum range
- Sentence position relative to the calculated guideline range (within, above, below)
- Plea type (guilty plea vs. trial)
- Fiscal year
Aggregation and Computed Metrics
After extraction, we aggregate case-level records by district and fiscal year to compute:
- Case counts: Total sentences per district per year per offense type.
- Average sentence: Mean imposed prison length in months; probation is coded as 0 months. Very long and life sentences are included at their recorded value in the USSC datafile.
- Guideline compliance rate: Percentage of cases with a determinable guideline range sentenced within that range. Cases with no calculable range are excluded from this share.
- Departure breakdown: Shares sentenced above or below the range (computed over cases with a determinable guideline range; a sentence outside the range is a departure or variance under 18 U.S.C. § 3553(a)).
- Disparity score: Average percentage difference vs. national average for the same offense types in each district, across all available years. A district at +20% sentences 20% longer than the national average for comparable offenses.
Corpus placement (#N of M)
Every published district and offense page states where that entity sits in PlainSentencing's own USSC-derived corpus. Placement is computed with SQL ROW_NUMBER over the same filtered sets used by the rankings - never hand-picked and never invented.
- District average sentence - rank among districts with at least 100 cases in the latest fiscal year, longest average first (same basis as /rankings/harshest/). Rank #1 is the harshest average sentence.
- District caseload - rank among districts with case_count > 0, most cases first (same basis as /rankings/most-cases/). Rank #1 is the busiest district.
- Offense caseload - rank among the 30 published offense types with case_count > 0, most federal defendants first.
- Offense average sentence - rank among those same offense types, longest average sentence first.
The Callout on each district and offense page names the applicable signals and links here so the derivation stays visible.
Important Coding Notes
- Life sentences are coded as 470 months in USSC data. This is included in average sentence calculations.
- Probation and other non-imprisonment sentences are coded as 0 months and are included in averages.
- No modifications are applied to the underlying USSC values.
Processing Pipeline
We download annual USSC public-use datafiles for each fiscal year and process them as follows:
- Parse fixed-width SPSS format files (FY2015–FY2023) and CSV format (FY2024–FY2025), extracting the relevant sentencing variables for each case record
- Map federal judicial district codes to district names and geographic locations
- Classify primary offense guideline categories using USSC's published offense type groupings
- Compute aggregate statistics (mean sentence, guideline compliance rate, departure breakdowns) by district, offense type, and fiscal year
- Calculate disparity scores comparing each district's sentencing patterns against national averages for the same offense types
No sentencing data is modified or editorialized. All case-level values (sentence length, guideline range, departure type) are taken directly from USSC public-use files. Aggregate metrics are computed from these source records using standard statistical methods.
Limitations
- Aggregate statistics cannot capture all case-specific factors that legitimately affect sentencing (quantity, role, cooperation, criminal history).
- Districts with lower case volumes show more volatile year-to-year changes.
- Disparity scores control for offense type but not for case-level aggravating or mitigating factors.
- Data is current through FY2025 (October 2024 – September 2025).
- USSC public-use files exclude certain identifying information for privacy; some case details are not available.
- The datafiles cover federal felony and Class A misdemeanor cases reported to the Commission by the district courts. Coverage is high but not universal: cases with incomplete sentencing documentation are excluded, so the figures reflect cases received and coded by the USSC and may modestly understate total federal sentencing activity.
- State criminal sentencing is entirely separate from federal sentencing and is not covered by this database.
Data reuse and access
The current reusable PlainSentencing dataset is the federal sentencing-by-district CSV. It contains the latest fiscal-year district-level aggregates shown on this site: district, state, circuit, case count, average sentence length, and the disparity measure. It is a derived aggregate of USSC public-use data, not a case-level data release.
The CSV is released under CC0 1.0. You may download, cite, quote, or republish it. Please identify PlainSentencing and the underlying USSC Individual Offender Datafiles when practical so readers can trace the figures to their source and methodology.
We do not currently sell licences, provide an API, offer a paid data product, or supply an embeddable chart widget. This page is a description of the access that exists today, not an offer to provide legal, commercial, or data services.
Revenue diversification readiness
This section documents what is productized today for reuse or licensing, how affiliates may appear if activated, and which alternative monetization paths are paperwork-ready. It is a readiness assessment only; nothing here activates new revenue without operator sign-off.
Downloadable and citable assets today
- District summary CSV at /data/plainsentencing-districts.csv - latest fiscal-year district aggregates (case count, average sentence, disparity). License: CC0 (also stated on /statistics).
- Flagship statistics page at /statistics - six Key Findings computed live from the USSC corpus, free to cite and link (CC0).
- Dataset JSON-LD on the homepage, methodology, statistics, about, privacy, terms, and selected hub pages declares the district extract with
isBasedOnpointing to USSC Individual Offender Datafiles. - CiteThis boxes on home, district and offense entity pages, rankings, statistics, methodology, guides, tools, and compare pages for one-click attribution of specific tables and counts.
- RSS feeds at
/feed.xml(site-wide) and per-entity/district/<slug>/feed.xml//offense/<slug>/feed.xmlfor FY vintage updates. - Machine discovery files at
/llms.txtand/llms-full.txtdescribing the USSC-derived corpus and primary page types for AI retrieval. - Per-entity OG images at district and offense share routes (1200×630 and 1600×900) generated from USSC aggregates, not hand-designed marketing art.
- Not yet productized: no bulk API endpoint, no embeddable chart iframe product, no paid data-licensing terms page. [operator] must approve pricing, API keys, and partnership terms before any launch.
Attribution ask: when reusing the CSV or citing a table, name PlainSentencing as the compiler and link to the source page or this methodology. CC0 dedicates the compiled aggregates; underlying USSC public-use files remain U.S. government works. Honest attribution supports reuse without implying USSC endorsement.
YMYL-safe affiliate placement rules (if activated)
PlainSentencing is a federal sentencing registry portal (POSTURE-B, H-YMYL legal). If the operator ever enables affiliate units beyond existing disclosure prose:
- Allowed page types: general federal-court process guides where the user is already reading educational context, never adjacent to a specific district or offense sentencing panel, departure breakdown, or caseload rank.
- Forbidden: legal advice, attorney referrals tied to a sentence outcome, or affiliate links on district detail pages, offense detail pages, rankings tables, compare pages, or any surface that could be read as predicting an individual sentence.
- Labeling: every paid link carries a visible affiliate disclosure; no disguised buttons inside data tables, guideline charts, or corpus-placement callouts.
- Density: same Module 36 ceilings as display ads; no affiliate block above the first USSC source-attribution line on registry pages.
Alternative ad-network eligibility (paperwork only)
- Google AdSense: portal status is
GETTING_READY(queued for review). Resubmit only after READY ≥90%, approval-gate green, and rendered-sibling-divergence exits the REJECT band (worst/offense/*class measured 86% on 2026-08-25). No manual ad slots are enabled pre-approval. - Mediavine Journey: requires ~1,000 real sessions per 28-day window. PlainSentencing measured ~68 Umami visitors and ~75 GA4 sessions in the latest 28-day window (2026-08-30), so it is not eligible today. Re-check when real sessions cross the threshold. [operator] applies to network enrollment.
- Data licensing / API: the CC0 district CSV + Dataset schema are the passive product surface; any commercial API or white-label embed requires [operator] pricing, terms of use, and legal review before publication.
Corrections and data changelog
PlainSentencing publishes descriptive statistics from USSC public-use files. If a figure on any page looks wrong, email hello@plainsentencing.com or use the contact form with (a) the exact page URL, (b) the field or sentence you believe is incorrect, and (c) a link to the official source you used to verify it.
We compare every report against the published USSC extract. Errors on our side are fixed in the database and ETL pipeline so every affected page updates together; if the page faithfully reflects the upstream file, we say so and point to the primary source. Material fixes are logged on our public data changelog; see also the editorial corrections process.
We aim to acknowledge data-error reports within 72 hours on business days. USSC annual releases (for example the FY2025 refresh on 2026-08-30) appear on the changelog when the pipeline ships; they are source updates, not corrections to prior mistakes.
Not Affiliated
PlainSentencing is not affiliated with the United States Sentencing Commission, the Department of Justice, or any government agency. Source data is public domain via ussc.gov.
Download the federal sentencing-by-district CSV cited on this page: plainsentencing-districts.csv (CC0).