Public record
Software health reportschema 0.31.0 · metrics 2.10.0 · 2026-08-05 04:05 UTC

huggingface / datasets

🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools

PythonApache-2.0★ 21,807 stars⑂ 3,333 forkssince Mar 2020View on GitHub ↗
KindCommand-line toolLibraryhow this is determined

huggingface/datasets holds a health index of 98 out of 100, placing it in the Exceptional band. It scores highest on Vitality (98/100) and lowest on AI Readiness (54/100). It was last updated 4 days ago. 2 contributors account for most of its recent work.

98
overall / 100
Exceptional

Software health index

Metrics are grouped into weighted categories on one standardized 1–100 scale. Overall starts as their weighted mean, calibrated against the distribution of the public record so bands carry percentile meaning; when public evidence triggers the High-Risk Jurisdiction Policy, the rating is adjusted and receives an At Risk ceiling of 34.

98
Exceptional93-100The record's top tier (≈ top 5%); essentially all checked criteria met
Excellent80-92Strong across the board; minor gaps
Good65-79Healthy; gaps are limited and manageable
Moderate50-64Acceptable with notable gaps; review recommended
Weak35-49Material weaknesses across several areas
At Risk20-34Significant weaknesses; adoption warrants caution
Critical1-19Severe problems (abandoned, single-maintainer, no hygiene)
VitalityCommunity &AdoptionSustainability &GovernanceEngineeringQualitySecurityAI Readiness

Score profile

Each axis is a category. The shape matters more than the average — a healthy subject fills the whole shape, while a spike-and-crater profile means strength in one dimension is masking risk in another.

The weighted overall 87 is calibrated to 98 on the published index scale (record calibration 2026-08-02).

Ownership

Hugging FaceOrganization
66,687 followers460 public repossince Feb 2017

This repository is backed by an organization — shared, accountable stewardship that can outlive any single maintainer.

Package ecosystems

RegistryPackageVersionDownloads / moVersionsLast publishTags
PyPIdatasets5.0.1145,695,8381197 days agodatasetsmachinelearning

Metrics by category

Vitality

Is the project alive — is code being written and are releases shipping?

98Exceptional · 21% of overall
How it's scored
36/36Push recencylast push 4 days ago
31.8/36Commit cadence46/52 weeks with commits
18/18Commit volume228 commits in the last year
10/10OpenSSF Scorecard: Maintained30 commit(s) and 4 issue activity found in the last 90 days -- score normalized to 10
Inputs used
commits_last_year228
human_commit_share1
days_since_last_push4
active_weeks_last_year46

Release discipline

100Exceptional
How it's scored
27/27Ships releases100 releases published
36/36Release recencylatest release 7 days ago
27/27Release cadencea release every ~16.7 days
0/10OpenSSF Scorecard: Signed-Releasesno data
Inputs used
releases_count100
latest_release_tag5.0.1
releases_from_tagsno
days_since_latest_release7
mean_days_between_releases16.7
Excluded from scoring (no data or not applicable): OpenSSF Scorecard: Signed-Releases. Remaining weights renormalized.

Community & Adoption

Does the project have users, downloads, attention, and a welcoming setup for contributors?

94Exceptional · 17% of overall
How it's scored
60/60Stars21,807 stars
25/25Forks3,333 forks
13.6/15Watchers279 watchers
Inputs used
forks3,333
stars21,807
watchers279
growth_stateunverified
growth_factor_pct100
growth_unverified_reasonno_history

Community health

85Excellent
How it's scored
22.5/22.5README
22.5/22.5Licenserecognized license (Apache-2.0)
18/18CONTRIBUTING guide
13.5/13.5Code of conduct
0/7.2Issue template
0/6.3PR template
Inputs used
has_readmeyes
has_licenseyes
readme_badges6
has_contributingyes
has_issue_templateno
has_code_of_conductyes
readme_badge_servicesgithub.com, shields.io
has_pull_request_templateno
How it's scored
80/80Monthly downloads145,695,838 downloads/month across pypi
0/20Registry dependentsnot reported by this ecosystem
Inputs used
packagesdatasets
dependents
ecosystemspypi
total_downloads
monthly_downloads145,695,838
unverified_packages_excluded
Excluded from scoring (no data or not applicable): Registry dependents. Remaining weights renormalized.

Sustainability & Governance

Will the project survive its people — bus factor, responsiveness, who backs it, and package upkeep?

83Excellent · 23% of overall
How it's scored
25.2/54Bus factor2 contributor(s) cover half of all commits
14.7/22.5Commit distributiontop contributor authored 35% of commits
13.5/13.5Contributor breadth100 contributors
10/10OpenSSF Scorecard: Contributorsproject has 30 contributing companies or organizations
Inputs used
bus_factor2
contributors_sampled100
top_contributor_share0.346
How it's scored
30.6/42Issue resolution73% of issues closed
25.6/30PR acceptance3,888/4,554 decided PRs merged
10.1/13Newcomer PR acceptance14/18 first-time contributors' PRs merged in 30d
10.5/15OpenSSF Scorecard: Code-ReviewFound 23/30 approved changesets -- score normalized to 7
Inputs used
merged_prs3,888
open_issues904
closed_issues2,430
prs_merged_7d1
prs_decided_7d3
prs_merged_30d27
prs_decided_30d33
issue_closed_ratio0.729
closed_unmerged_prs666
first_time_authors_30d12
first_time_prs_merged_30d14
first_time_prs_decided_30d18
How it's scored
30/30Ownership backingorganization-owned
0/20Verified domainverified-domain status not read for this organization
25/25Owner reach66,687 followers of huggingface
25/25Track record460 public repos, account ~9 yr old
Inputs used
followers66,687
owner_typeOrganization
is_verified
owner_loginhuggingface
public_repos460
account_age_days3,460
Excluded from scoring (no data or not applicable): Verified domain. Remaining weights renormalized.

Package maintenance

100Exceptional
How it's scored
25/25Published & resolvable1 package(s) on pypi
35/35Publish recencylatest publish 7 days ago
20/20Version history119 published versions
20/20Not deprecatedactive, not deprecated or yanked
Inputs used
packagesdatasets
ecosystemspypi
any_deprecatedno
min_days_since_publish7

Engineering Quality

Are baseline engineering and documentation practices in place?

95Exceptional · 19% of overall
How it's scored
24/24CI workflows7 workflow(s)
24/24Tests present
16/16Linter config
9.6/9.6Pre-commit hooks
0/6.4.editorconfig
18/20OpenSSF Scorecard: CI-Tests28 out of 30 merged PRs checked by a CI test -- score normalized to 9
Inputs used
has_ciyes
has_testsyes
has_editorconfigno
has_linter_configyes
has_precommit_configyes

Documentation

100Exceptional
How it's scored
30/30README
25/25Documentation directory
15/15Documentation / homepage sitehttps://huggingface.co/docs/datasets
10/10Repository description
10/10Topics16 topics
10/10Wiki
Inputs used
topicsnlp, datasets, pytorch, tensorflow, pandas, numpy, natural-language-processing, computer-vision, machine-learning, deep-learning, speech, ai, artificial-intelligence, llm, dataset-hub, huggingface
has_wikiyes
homepagehttps://huggingface.co/docs/datasets
docs_sitehttps://huggingface.co/docs/datasets
has_readmeyes
has_docs_diryes
has_descriptionyes

Security

Are visible security and supply-chain practices strong, without unresolved high-risk jurisdiction exposure?

70Good · 16% of overall
How it's scored
7.5/7.5Binary-Artifactsno binaries found in the repo
0/7.5Branch-Protectionno data
2.2/2.5CI-Tests28 out of 30 merged PRs checked by a CI test -- score normalized to 9
0/2.5CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
5.2/7.5Code-ReviewFound 23/30 approved changesets -- score normalized to 7
2.5/2.5Contributorsproject has 30 contributing companies or organizations
10/10Dangerous-Workflowno dangerous workflow patterns detected
0/7.5Dependency-Update-Toolno update tool detected
0/5Fuzzingproject is not fuzzed
2.5/2.5Licenselicense file detected
7.5/7.5Maintained30 commit(s) and 4 issue activity found in the last 90 days -- score normalized to 10
0/5Packagingno data
2.5/5Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 5
0/5SASTSAST tool is not run on all commits -- score normalized to 0
5/5Security-Policysecurity policy file detected
0/7.5Signed-Releasesno data
0/7.5Token-Permissionsdetected GitHub workflow tokens with excessive permissions
7.5/7.5Vulnerabilities0 existing vulnerabilities detected
Inputs used
sourceopenssf_scorecard
checks_evaluated15
scorecard_versionv5.5.0
checks_inconclusive3
scorecard_aggregate6.2
Excluded from scoring (no data or not applicable): Branch-Protection, Packaging, Signed-Releases. Remaining weights renormalized.

Dependency advisories

100Exceptional
How it's scored
35/35Direct dependencies free of known advisoriesno direct dependency carries a known advisory
25/25Indirect dependencies free of known advisoriesno indirect dependency carries a known advisory
0/40No advisories left outstandingno advisory carries a publication date
Inputs used
sourceosv
advisories0
affected_packages0
assessed_packages34
unassessed_packages0
affected_by_severitynone
direct_affected_packages0
Excluded from scoring (no data or not applicable): No advisories left outstanding. Remaining weights renormalized. Matched the pypi:datasets@5.0.1 runtime dependency closure — what installing the published package pulls in — 34 packages. Reachability is not analyzed.

AI Readiness

How well is the repo equipped to be developed and maintained with AI coding agents? Carries a deliberately small weight (4%): agent tooling is a real maintenance signal, but a repository with none can still reach 100/100.

54Moderate · 4% of overall
How it's scored
0/45Agent instructionsno CLAUDE.md / AGENTS.md / editor rules
0/15Machine-readable docs (llms.txt)
40/40Legible commit history100 of 100 human commits state their intent (structured subject or explanatory body)
Inputs used
has_llms_txtno
llms_txt_url
legible_history_share1
agent_instruction_files
agent_instruction_max_bytes
How it's scored
18/18One-command bootstrapMakefile
22/22Automated tests
11/11Lint / format config
0/11Static type checking
0/10Reproducible environment
10/10Demonstrated agent practice7 of the last 100 commits agent-authored or agent-credited
0/8Automated maintenanceno automated dependency updates observed
5/10OpenSSF Scorecard: Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 5
Inputs used
has_nixno
has_testsyes
lockfiles
has_dockerfileno
typed_languageno
bootstrap_filesMakefile
has_devcontainerno
has_linter_configyes
typecheck_configs
agent_commit_share0.07
toolchain_manifests
dependency_bot_commit_share0
How it's scored
0/45Type-checkable codePython without a type-check config
52.4/55Manageable file sizes11/234 source files over 60KB
Inputs used
primary_languagePython
largest_source_bytes343,664
source_files_sampled234
oversized_source_files11

Key facts

21,807GitHub stars
100contributors
228commits, last 12 months
4days since last push
100releases
2bus factor
904open issues
PyPIpackage ecosystems

Data collection warnings

  • Star history unavailable: GitHub GraphQL error: Resource not accessible by personal access token
  • First-time contributor figures cover 12 of 18 authors (cap 12)

More detail

Star and fork history 0 ★ / 3,333 ⇿
0Stars
3,333Forks
39Releases

When each star and fork was added, collected from GitHub and bucketed by day. Cumulative growth sits directly above the daily additions it is made of, so the two read against each other: steady organic accretion looks nothing like an abrupt, short-lived burst. Where that difference is measurable, it is reported as growth authenticity.

Only the most recent history is shown — this repository exceeds the collection window, so the earliest history is not captured.

2,2002,4002,6002,8003,0003,2003,4003,333102024-022025-052026-08
Major 3Minor 18Patch 18

Each point covers 3 days.

OpenSSF Scorecard 6.2 / 10
6.2aggregate

Independent, tool-agnostic security assessment from the open-source OpenSSF Scorecard. Each check rewards a security practice, not a specific vendor's tool. Checks Scorecard could not determine are marked n/a and excluded from the security score (never counted as zero).Scorecard v5.5.0 · 2026-08-05 04:05 UTC

10Binary-Artifactsno binaries found in the repo
n/aBranch-Protectioninternal error: error during branchesHandler.setup: internal error: some github tokens can't read classic branch protection rules: https://github.com/ossf/scorecard-action/blob/main/docs/authentication/fine-grained-auth-token.md
9CI-Tests28 out of 30 merged PRs checked by a CI test -- score normalized to 9
0CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
7Code-ReviewFound 23/30 approved changesets -- score normalized to 7
10Contributorsproject has 30 contributing companies or organizations
10Dangerous-Workflowno dangerous workflow patterns detected
0Dependency-Update-Toolno update tool detected
0Fuzzingproject is not fuzzed
10Licenselicense file detected
10Maintained30 commit(s) and 4 issue activity found in the last 90 days -- score normalized to 10
n/aPackagingpackaging workflow not detected
5Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 5
0SASTSAST tool is not run on all commits -- score normalized to 0
10Security-Policysecurity policy file detected
n/aSigned-Releasesno releases found
0Token-Permissionsdetected GitHub workflow tokens with excessive permissions
10Vulnerabilities0 existing vulnerabilities detected
All dependencies 18

Full resolved dependency set from the GitHub dependency graph: 0 direct and 18 indirect (transitive) packages. The transitive closure is complete when the repository commits a lockfile.

RegistryPackageVersionRelation
PyPIdillindirect
PyPIfilelockindirect
PyPIfsspecindirect
PyPIgitpython3.1.27indirect
PyPIhttpxindirect
PyPIhuggingface-hubindirect
PyPImultiprocessindirect
PyPInumpyindirect
PyPIpackagingindirect
PyPIpandasindirect
PyPIpyarrowindirect
PyPIpython-dotenv0.19.2indirect
PyPIpyyamlindirect
PyPIrequestsindirect
PyPIrequests2.25.1indirect
PyPItqdmindirect
PyPItqdm4.62.3indirect
PyPIxxhashindirect
Dependency advisories 0

Installing pypi:datasets@5.0.1 pulls in 34 packages, direct and transitive: 0 carry known advisories, of which 0 are direct dependencies.

No known advisories affect the assessed dependencies.

An advisory means the version recorded in the dependency graph falls inside an advisory’s affected range. Reachability is not analysed, and the graph includes development and test pins — a finding may concern tooling rather than shipped software.

Raw JSON report machine-readable

Feedback

Spotted something off in this report, or have thoughts to share? Wrong measurements, missed tooling, ideas, questions — anything is welcome. Every message is read and gets a response.

The message is kept through sign-in.

Scores are signals, not warranties. They reflect publicly visible practices on GitHub — not a code audit, and not a security guarantee.

Missing data is excluded and weights renormalized, never scored as zero. Methodology is versioned and open: metrics v2.10.0, schema v0.31.0 — full methodology · metrics wiki.

How one result sits in the wider record: aggregate statisticsPyPI.