Public record
Software health reportschema 0.31.0 · metrics 2.10.0 · 2026-08-13 03:17 UTC

pemistahl / lingua-py

The most accurate natural language detection library for Python, suitable for short text and mixed-language text

PythonApache-2.0★ 1,775 stars⑂ 61 forkssince Jul 2021View on GitHub ↗

pemistahl/lingua-py holds a health index of 51 out of 100, placing it in the Moderate band. It scores highest on Community & Adoption (62/100) and lowest on AI Readiness (22/100). It was last updated 23 days ago. A single contributor accounts for most of its recent work.

51
overall / 100
Moderate

Software health index

Metrics are grouped into weighted categories on one standardized 1–100 scale. Overall starts as their weighted mean, calibrated against the distribution of the public record so bands carry percentile meaning; when public evidence triggers the High-Risk Jurisdiction Policy, the rating is adjusted and receives an At Risk ceiling of 34.

51
Exceptional93-100The record's top tier (≈ top 5%); essentially all checked criteria met
Excellent80-92Strong across the board; minor gaps
Good65-79Healthy; gaps are limited and manageable
Moderate50-64Acceptable with notable gaps; review recommended
Weak35-49Material weaknesses across several areas
At Risk20-34Significant weaknesses; adoption warrants caution
Critical1-19Severe problems (abandoned, single-maintainer, no hygiene)
VitalityCommunity &AdoptionSustainability &GovernanceEngineeringQualitySecurityAI Readiness

Score profile

Each axis is a category. The shape matters more than the average — a healthy subject fills the whole shape, while a spike-and-crater profile means strength in one dimension is masking risk in another.

The weighted overall 51 is calibrated to 51 on the published index scale (record calibration 2026-08-02).

Ownership

Peter M. StahlPersonal account
296 followers15 public repossince Oct 2011@riege

This repository is owned by a personal account. A single-owner project carries more continuity risk than an organization-backed one.

Package ecosystems

RegistryPackageVersionDownloads / moVersionsLast publishTags
PyPIlingua-language-detector2.2.0-21156 days agolanguage-processinglanguage-detectionlanguage-recognitionnlp

Metrics by category

Vitality

Is the project alive — is code being written and are releases shipping?

55Moderate · 21% of overall
How it's scored
28.8/36Push recencylast push 23 days ago
2.8/36Commit cadence4/52 weeks with commits
10.6/18Commit volume14 commits in the last year
0/10OpenSSF Scorecard: Maintained0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0
Inputs used
commits_last_year14
human_commit_share0.79
days_since_last_push23
active_weeks_last_year4
How it's scored
27/27Ships releases23 releases published
27/36Release recencylatest release 156 days ago
19.8/27Release cadencea release every ~93.9 days
0/10OpenSSF Scorecard: Signed-ReleasesProject has not signed or included provenance with any releases.
Inputs used
releases_count23
latest_release_tagv2.2.0
releases_from_tagsno
days_since_latest_release156
mean_days_between_releases93.9

Community & Adoption

Does the project have users, downloads, attention, and a welcoming setup for contributors?

62Moderate · 17% of overall
How it's scored
52.7/60Stars1,775 stars
14.8/25Forks61 forks
5.3/15Watchers10 watchers
Inputs used
forks61
stars1,775
watchers10
growth_stateunverified
growth_factor_pct100
growth_unverified_reasonno_history
How it's scored
22.5/22.5README
22.5/22.5Licenserecognized license (Apache-2.0)
0/18CONTRIBUTING guide
0/13.5Code of conduct
0/7.2Issue template
0/6.3PR template
Inputs used
has_readmeyes
has_licenseyes
readme_badges6
has_contributingno
has_issue_templateno
has_code_of_conductno
readme_badge_servicescodecov.io, github.com, shields.io
has_pull_request_templateno

Sustainability & Governance

Will the project survive its people — bus factor, responsiveness, who backs it, and package upkeep?

55Moderate · 23% of overall
How it's scored
9/54Bus factor1 contributor(s) cover half of all commits
1/22.5Commit distributiontop contributor authored 96% of commits
9.5/13.5Contributor breadth7 contributors
3/10OpenSSF Scorecard: Contributorsproject has 1 contributing companies or organizations -- score normalized to 3
Inputs used
bus_factor1
contributors_sampled7
top_contributor_share0.957
How it's scored
28.8/42Issue resolution68% of issues closed
15.1/30PR acceptance72/143 decided PRs merged
0/13Newcomer PR acceptanceno first-time contributor's PR decided in 30d
1.5/15OpenSSF Scorecard: Code-ReviewFound 3/20 approved changesets -- score normalized to 1
Inputs used
merged_prs72
open_issues39
closed_issues85
prs_merged_7d0
prs_decided_7d0
prs_merged_30d0
prs_decided_30d0
issue_closed_ratio0.685
closed_unmerged_prs71
first_time_authors_30d0
first_time_prs_merged_30d0
first_time_prs_decided_30d0
Excluded from scoring (no data or not applicable): Newcomer PR acceptance. Remaining weights renormalized.
How it's scored
10/30Ownership backingpersonal (user) account
0/20Verified domainnot applicable to user accounts
17.8/25Owner reach296 followers of pemistahl
20.8/25Track record15 public repos, account ~14 yr old
Inputs used
followers296
owner_typeUser
is_verified
owner_loginpemistahl
public_repos15
account_age_days5,408
Excluded from scoring (no data or not applicable): Verified domain. Remaining weights renormalized.

Package maintenance

100Exceptional
How it's scored
25/25Published & resolvable1 package(s) on pypi
35/35Publish recencylatest publish 156 days ago
20/20Version history21 published versions
20/20Not deprecatedactive, not deprecated or yanked
Inputs used
packageslingua-language-detector
ecosystemspypi
any_deprecatedno
min_days_since_publish156

Engineering Quality

Are baseline engineering and documentation practices in place?

42Weak · 19% of overall
How it's scored
24/24CI workflows1 workflow(s)
0/24Tests present
0/16Linter config
0/9.6Pre-commit hooks
6.4/6.4.editorconfig
6/20OpenSSF Scorecard: CI-Tests4 out of 13 merged PRs checked by a CI test -- score normalized to 3
Inputs used
has_ciyes
has_testsno
has_editorconfigyes
has_linter_configno
has_precommit_configno

Documentation

50Moderate
How it's scored
30/30README
0/25Documentation directory
0/15Documentation / homepage site
10/10Repository description
10/10Topics7 topics
0/10Wiki
Inputs used
topicsnlp, natural-language-processing, language-detection, language-recognition, language-identification, language-classification, python-library
has_wikino
homepage
docs_site
has_readmeyes
has_docs_dirno
has_descriptionyes

Security

Are visible security and supply-chain practices strong, without unresolved high-risk jurisdiction exposure?

44Weak · 16% of overall
How it's scored
7.5/7.5Binary-Artifactsno binaries found in the repo
0/7.5Branch-Protectionbranch protection not enabled on development/release branches
0.8/2.5CI-Tests4 out of 13 merged PRs checked by a CI test -- score normalized to 3
0/2.5CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
0.8/7.5Code-ReviewFound 3/20 approved changesets -- score normalized to 1
0.8/2.5Contributorsproject has 1 contributing companies or organizations -- score normalized to 3
10/10Dangerous-Workflowno dangerous workflow patterns detected
7.5/7.5Dependency-Update-Toolupdate tool detected
0/5Fuzzingproject is not fuzzed
2.5/2.5Licenselicense file detected
0/7.5Maintained0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0
0/5Packagingno data
0/5Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
0/5SASTSAST tool is not run on all commits -- score normalized to 0
0/5Security-Policysecurity policy file not detected
0/7.5Signed-ReleasesProject has not signed or included provenance with any releases.
0/7.5Token-Permissionsdetected GitHub workflow tokens with excessive permissions
0/7.5Vulnerabilities21 existing vulnerabilities detected
Inputs used
sourceopenssf_scorecard
checks_evaluated17
scorecard_versionv5.5.0
checks_inconclusive1
scorecard_aggregate3
Excluded from scoring (no data or not applicable): Packaging. Remaining weights renormalized.

Dependency advisories

100Exceptional
How it's scored
35/35Direct dependencies free of known advisoriesno direct dependency carries a known advisory
0/25Indirect dependencies free of known advisoriestransitive set not separable from development and test dependencies in this scope
0/40No advisories left outstandingno advisory carries a publication date
Inputs used
sourceosv
advisories40
affected_packages3
assessed_packages31
unassessed_packages0
affected_by_severitycritical 2, high 1
direct_affected_packages0
Excluded from scoring (no data or not applicable): Indirect dependencies free of known advisories, No advisories left outstanding. Remaining weights renormalized. Matched 31 resolved dependencies against OSV. This repository publishes no package the index resolves, so the repository dependency graph was assessed instead. That graph mixes development and test pins with shipped dependencies, so only the declared runtime dependencies are scored; transitive findings are reported as context and excluded from the score. Reachability is not analyzed.

AI Readiness

How well is the repo equipped to be developed and maintained with AI coding agents? Carries a deliberately small weight (4%): agent tooling is a real maintenance signal, but a repository with none can still reach 100/100.

22At Risk · 4% of overall
How it's scored
0/45Agent instructionsno CLAUDE.md / AGENTS.md / editor rules
0/15Machine-readable docs (llms.txt)
10.1/40Legible commit history15 of 79 human commits state their intent (structured subject or explanatory body)
Inputs used
has_llms_txtno
llms_txt_url
legible_history_share0.19
agent_instruction_files
agent_instruction_max_bytes
How it's scored
0/18One-command bootstrap
0/22Automated tests
0/11Lint / format config
0/11Static type checking
10/10Reproducible environmentlockfile
0/10Demonstrated agent practiceno agent-authored commits among the last 100
8/8Automated maintenance21 of the last 100 commits are automated dependency updates
0/10OpenSSF Scorecard: Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
Inputs used
has_nixno
has_testsno
lockfilespoetry.lock
has_dockerfileno
typed_languageno
bootstrap_files
has_devcontainerno
has_linter_configno
typecheck_configs
agent_commit_share0
toolchain_manifests
dependency_bot_commit_share0.21
How it's scored
0/45Type-checkable codePython without a type-check config
55/55Manageable file sizes0/3 source files over 60KB
Inputs used
primary_languagePython
largest_source_bytes21,595
source_files_sampled3
oversized_source_files0

Key facts

1,775GitHub stars
7contributors
14commits, last 12 months
23days since last push
23releases
1bus factor
39open issues
PyPIpackage ecosystems

Data collection warnings

  • Star history unavailable: GitHub GraphQL error: Resource not accessible by personal access token

More detail

Star and fork history 0 ★ / 61 ⇿
0Stars
61Forks
23Releases

When each star and fork was added, collected from GitHub and bucketed by day. Cumulative growth sits directly above the daily additions it is made of, so the two read against each other: steady organic accretion looks nothing like an abrupt, short-lived burst. Where that difference is measurable, it is reported as growth authenticity.

013253850636122022-012024-032026-06
Major 2Minor 6Patch 15

Each point covers 5 days.

OpenSSF Scorecard 3.0 / 10
3.0aggregate

Independent, tool-agnostic security assessment from the open-source OpenSSF Scorecard. Each check rewards a security practice, not a specific vendor's tool. Checks Scorecard could not determine are marked n/a and excluded from the security score (never counted as zero).Scorecard v5.5.0 · 2026-08-13 03:16 UTC

10Binary-Artifactsno binaries found in the repo
0Branch-Protectionbranch protection not enabled on development/release branches
3CI-Tests4 out of 13 merged PRs checked by a CI test -- score normalized to 3
0CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
1Code-ReviewFound 3/20 approved changesets -- score normalized to 1
3Contributorsproject has 1 contributing companies or organizations -- score normalized to 3
10Dangerous-Workflowno dangerous workflow patterns detected
10Dependency-Update-Toolupdate tool detected
0Fuzzingproject is not fuzzed
10Licenselicense file detected
0Maintained0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0
n/aPackagingpackaging workflow not detected
0Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
0SASTSAST tool is not run on all commits -- score normalized to 0
0Security-Policysecurity policy file not detected
0Signed-ReleasesProject has not signed or included provenance with any releases.
0Token-Permissionsdetected GitHub workflow tokens with excessive permissions
0Vulnerabilities21 existing vulnerabilities detected
All dependencies 31

Full resolved dependency set from the GitHub dependency graph: 0 direct and 31 indirect (transitive) packages. The transitive closure is complete when the repository commits a lockfile.

RegistryPackageVersionRelation
PyPIblack25.12.0indirect
PyPIclick8.3.1indirect
PyPIcolorama0.4.6indirect
PyPIcontourpy1.3.3indirect
PyPIcycler0.12.1indirect
PyPIfonttools4.61.1indirect
PyPIgcld33.0.13indirect
PyPIkiwisolver1.4.9indirect
PyPIlangdetect1.0.9indirect
PyPIlangid1.1.6indirect
PyPIlibrt0.8.0indirect
PyPImatplotlib3.10.8indirect
PyPImypy1.19.1indirect
PyPImypy-extensions1.1.0indirect
PyPInumpy2.4.2indirect
PyPIpackaging26.0indirect
PyPIpandas2.3.3indirect
PyPIpathspec1.0.4indirect
PyPIpillow12.1.1indirect
PyPIplatformdirs4.7.0indirect
PyPIpsutil7.0.0indirect
PyPIpycld20.42indirect
PyPIpyparsing3.3.2indirect
PyPIpython-dateutil2.9.0.post0indirect
PyPIpytokens0.4.1indirect
PyPIpytz2025.2indirect
PyPIseaborn0.13.2indirect
PyPIsimplemma0.9.1indirect
PyPIsix1.17.0indirect
PyPItyping-extensions4.15.0indirect
PyPItzdata2025.3indirect
Dependency advisories 3

This repository publishes no package the index resolves, so its own dependency graph was assessed — 31 packages, which also include development and test pins that never ship: 3 carry known advisories, of which 0 are direct.

PackageVersionRelationSeverityAdvisoriesFixed in
black25.12.0indirectcritical326.3.1
pillow12.1.1indirectcritical3612.3.0
click8.3.1indirecthigh18.3.3

An advisory means the version recorded in the dependency graph falls inside an advisory’s affected range. Reachability is not analysed, and the graph includes development and test pins — a finding may concern tooling rather than shipped software.

Raw JSON report machine-readable

Feedback

Spotted something off in this report, or have thoughts to share? Wrong measurements, missed tooling, ideas, questions — anything is welcome. Every message is read and gets a response.

The message is kept through sign-in.

Scores are signals, not warranties. They reflect publicly visible practices on GitHub — not a code audit, and not a security guarantee.

Missing data is excluded and weights renormalized, never scored as zero. Methodology is versioned and open: metrics v2.10.0, schema v0.31.0 — full methodology · metrics wiki.

How one result sits in the wider record: aggregate statisticsPyPI.