Public record
Software health reportschema 0.31.0 · metrics 2.5.0 · 2026-08-08 19:37 UTC

PyThaiNLP / pythainlp

Thai natural language processing in Python

PythonApache-2.0★ 1,148 stars⑂ 300 forkssince Jun 2016View on GitHub ↗
KindCommand-line toolLibraryhow this is determined

PyThaiNLP/pythainlp holds a health index of 98 out of 100, placing it in the Exceptional band. It scores highest on Engineering Quality (94/100) and lowest on AI Readiness (76/100). It was last updated today. 2 contributors account for most of its recent work.

98
overall / 100
Exceptional

Software health index

Metrics are grouped into weighted categories on one standardized 1–100 scale. Overall starts as their weighted mean, calibrated against the distribution of the public record so bands carry percentile meaning; when public evidence triggers the High-Risk Jurisdiction Policy, the rating is adjusted and receives an At Risk ceiling of 34.

98
Exceptional93-100The record's top tier (≈ top 5%); essentially all checked criteria met
Excellent80-92Strong across the board; minor gaps
Good65-79Healthy; gaps are limited and manageable
Moderate50-64Acceptable with notable gaps; review recommended
Weak35-49Material weaknesses across several areas
At Risk20-34Significant weaknesses; adoption warrants caution
Critical1-19Severe problems (abandoned, single-maintainer, no hygiene)
VitalityCommunity &AdoptionSustainability &GovernanceEngineeringQualitySecurityAI Readiness

Score profile

Each axis is a category. The shape matters more than the average — a healthy subject fills the whole shape, while a spike-and-crater profile means strength in one dimension is masking risk in another.

The weighted overall 86 is calibrated to 98 on the published index scale (record calibration 2026-08-02).

Ownership

PyThaiNLPOrganization
182 followers111 public repossince Oct 2017

This repository is backed by an organization — shared, accountable stewardship that can outlive any single maintainer.

Package ecosystems

Metrics by category

Vitality

Is the project alive — is code being written and are releases shipping?

93Exceptional · 21% of overall
How it's scored
36/36Push recency — last push 0 days ago
29.8/36Commit cadence — 43/52 weeks with commits
18/18Commit volume — 1,520 commits in the last year
10/10OpenSSF Scorecard: Maintained — 30 commit(s) and 7 issue activity found in the last 90 days -- score normalized to 10
Inputs used
commits_last_year1,520
human_commit_share0.82
days_since_last_push0
active_weeks_last_year43
How it's scored
27/27Ships releases — 100 releases published
36/36Release recency — latest release 10 days ago
19.8/27Release cadence — a release every ~53.9 days
0/10OpenSSF Scorecard: Signed-Releases — no data
Inputs used
releases_count100
latest_release_tagv5.3.5
releases_from_tagsno
days_since_latest_release10
mean_days_between_releases53.9
Excluded from scoring (no data or not applicable): OpenSSF Scorecard: Signed-Releases. Remaining weights renormalized.

Community & Adoption

Does the project have users, downloads, attention, and a welcoming setup for contributors?

89Excellent · 17% of overall
How it's scored
49.6/60Stars — 1,148 stars
20.6/25Forks — 300 forks
9/15Watchers — 42 watchers
Inputs used
forks300
stars1,148
watchers42
growth_stateunverified
growth_factor_pct100
growth_unverified_reasonno_history

Community health

92Excellent
How it's scored
22.5/22.5README
22.5/22.5License — recognized license (Apache-2.0)
18/18CONTRIBUTING guide
13.5/13.5Code of conduct
0/7.2Issue template
6.3/6.3PR template
Inputs used
has_readmeyes
has_licenseyes
readme_badges7
has_contributingyes
has_issue_templateno
has_code_of_conductyes
readme_badge_servicesapp.codacy.com, badgen.net, coveralls.io, shields.io
has_pull_request_templateyes
How it's scored
80/80Monthly downloads — 1,510,408 downloads/month across pypi
0/20Registry dependents — not reported by this ecosystem
Inputs used
packagespythainlp
dependents
ecosystemspypi
total_downloads
monthly_downloads1,510,408
Excluded from scoring (no data or not applicable): Registry dependents. Remaining weights renormalized.

Sustainability & Governance

Will the project survive its people — bus factor, responsiveness, who backs it, and package upkeep?

78Good · 23% of overall
How it's scored
25.2/54Bus factor — 2 contributor(s) cover half of all commits
11.5/22.5Commit distribution — top contributor authored 49% of commits
13.5/13.5Contributor breadth — 60 contributors
10/10OpenSSF Scorecard: Contributors — project has 21 contributing companies or organizations
Inputs used
bus_factor2
contributors_sampled60
top_contributor_share0.488
How it's scored
39.8/42Issue resolution — 95% of issues closed
26.5/30PR acceptance — 886/1,002 decided PRs merged
13/13Newcomer PR acceptance — 2/2 first-time contributors' PRs merged in 30d
7.5/15OpenSSF Scorecard: Code-Review — Found 3/6 approved changesets -- score normalized to 5
Inputs used
merged_prs886
open_issues23
closed_issues422
prs_merged_7d2
prs_decided_7d2
prs_merged_30d6
prs_decided_30d6
issue_closed_ratio0.948
closed_unmerged_prs116
first_time_authors_30d2
first_time_prs_merged_30d2
first_time_prs_decided_30d2
How it's scored
30/30Ownership backing — organization-owned
0/20Verified domain
16.3/25Owner reach — 182 followers of PyThaiNLP
25/25Track record — 111 public repos, account ~8 yr old
Inputs used
followers182
owner_typeOrganization
is_verified
owner_loginPyThaiNLP
public_repos111
account_age_days3,215

Package maintenance

100Exceptional
How it's scored
25/25Published & resolvable — 1 package(s) on pypi
35/35Publish recency — latest publish 10 days ago
20/20Version history — 119 published versions
20/20Not deprecated — active, not deprecated or yanked
Inputs used
packagespythainlp
ecosystemspypi
any_deprecatedno
min_days_since_publish10

Engineering Quality

Are baseline engineering and documentation practices in place?

94Exceptional · 19% of overall
How it's scored
24/24CI workflows — 14 workflow(s)
24/24Tests present
16/16Linter config — .flake8
0/9.6Pre-commit hooks
6.4/6.4.editorconfig
20/20OpenSSF Scorecard: CI-Tests — 11 out of 11 merged PRs checked by a CI test -- score normalized to 10
Inputs used
has_ciyes
has_testsyes
has_editorconfigyes
has_linter_configyes
has_precommit_configno

Documentation

100Exceptional
How it's scored
30/30README
25/25Documentation directory
15/15Documentation / homepage site — https://pythainlp.org/
10/10Repository description
10/10Topics — 14 topics
10/10Wiki
Inputs used
topicspython, thai-nlp, nlp-library, thai-language, natural-language-processing, thai-nlp-library, thai-soundex, soundex, word-segmentation, thai, hacktoberfest, computational-linguistics, text-processing, hacktoberfest-accepted
has_wikiyes
homepagehttps://pythainlp.org/
has_readmeyes
has_docs_diryes
has_descriptionyes

Security

Are visible security and supply-chain practices strong, without unresolved high-risk jurisdiction exposure?

79Good · 16% of overall
How it's scored
7.5/7.5Binary-Artifacts — no binaries found in the repo
0/7.5Branch-Protection — branch protection not enabled on development/release branches
2.5/2.5CI-Tests — 11 out of 11 merged PRs checked by a CI test -- score normalized to 10
0.5/2.5CII-Best-Practices — badge detected: InProgress
3.8/7.5Code-Review — Found 3/6 approved changesets -- score normalized to 5
2.5/2.5Contributors — project has 21 contributing companies or organizations
10/10Dangerous-Workflow — no dangerous workflow patterns detected
7.5/7.5Dependency-Update-Tool — update tool detected
5/5Fuzzing — project is fuzzed
2.5/2.5License — license file detected
7.5/7.5Maintained — 30 commit(s) and 7 issue activity found in the last 90 days -- score normalized to 10
5/5Packaging — packaging workflow detected
0/5Pinned-Dependencies — dependency not pinned by hash detected -- score normalized to 0
5/5SAST — SAST tool is run on all commits
5/5Security-Policy — security policy file detected
0/7.5Signed-Releases — no data
0/7.5Token-Permissions — detected GitHub workflow tokens with excessive permissions
7.5/7.5Vulnerabilities — 0 existing vulnerabilities detected
Inputs used
sourceopenssf_scorecard
checks_evaluated17
scorecard_versionv5.5.0
checks_inconclusive1
scorecard_aggregate7.4
Excluded from scoring (no data or not applicable): signed_releases. Remaining weights renormalized.

Dependency advisories

100Exceptional
How it's scored
35/35Direct dependencies free of known advisories — no direct dependency carries a known advisory
0/25Indirect dependencies free of known advisories — transitive set not separable from development and test dependencies in this scope
0/40No advisories left outstanding — no advisory carries a publication date
Inputs used
sourceosv
advisories0
affected_packages0
assessed_packages19
unassessed_packages48
affected_by_severitynone
direct_affected_packages0
Excluded from scoring (no data or not applicable): Indirect dependencies free of known advisories, No advisories left outstanding. Remaining weights renormalized. Matched 19 resolved dependencies against OSV. 48 could not be assessed — no resolved version, an unsupported ecosystem, or beyond the reported package list. This repository publishes no package the index resolves, so the repository dependency graph was assessed instead. That graph mixes development and test pins with shipped dependencies, so only the declared runtime dependencies are scored; transitive findings are reported as context and excluded from the score. Reachability is not analyzed.

AI Readiness

How well is the repo equipped to be developed and maintained with AI coding agents? Carries a deliberately small weight (4%): agent tooling is a real maintenance signal, but a repository with none can still reach 100/100.

76Good · 4% of overall
How it's scored
45/45Agent instructions — .github/copilot-instructions.md, AGENTS.md
0/15Machine-readable docs (llms.txt)
36.4/40Legible commit history — 56 of 82 human commits state their intent (structured subject or explanatory body)
Inputs used
has_llms_txtno
legible_history_share0.683
agent_instruction_files.github/copilot-instructions.md, AGENTS.md
agent_instruction_max_bytes19,770
How it's scored
18/18One-command bootstrap — Makefile, docs/Makefile
22/22Automated tests
11/11Lint / format config — .flake8
11/11Static type checking — pythainlp/py.typed
10/10Reproducible environment — Dockerfile
4/10Demonstrated agent practice — 2 of the last 100 commits agent-authored or agent-credited
8/8Automated maintenance — 18 of the last 100 commits are automated dependency updates
0/10OpenSSF Scorecard: Pinned-Dependencies — dependency not pinned by hash detected -- score normalized to 0
Inputs used
has_nixno
has_testsyes
lockfiles
has_dockerfileyes
typed_languageno
bootstrap_filesMakefile, docs/Makefile
has_devcontainerno
has_linter_configyes
typecheck_configspythainlp/py.typed
agent_commit_share0.02
toolchain_manifests
dependency_bot_commit_share0.18
How it's scored
27/45Type-checkable code — Python with type-check config (pythainlp/py.typed)
54.8/55Manageable file sizes — 1/290 source files over 60KB
Inputs used
primary_languagePython
largest_source_bytes93,583
source_files_sampled290
oversized_source_files1
How it's scored
0/40API schema (OpenAPI/GraphQL/proto)
0/20MCP server
40/40Runnable examples — examples, notebooks
Inputs used
example_dirsexamples, notebooks
has_mcp_signalno
api_schema_files

Key facts

1,148GitHub stars
60contributors
1,520commits, last 12 months
0days since last push
100releases
2bus factor
23open issues
PyPIpackage ecosystems

Data collection warnings

  • Star history unavailable: GitHub GraphQL error: Resource not accessible by personal access token

More detail

Star and fork history 0 ★ / 300 ⇿
0Stars
300Forks
99Releases

When each star and fork was added, collected from GitHub and bucketed by day. Cumulative growth sits directly above the daily additions it is made of, so the two read against each other: steady organic accretion looks nothing like an abrupt, short-lived burst. Where that difference is measurable, it is reported as growth authenticity.

05010015020025030029882016-062021-072026-07
Major 3Minor 7Patch 47

Each point covers 10 days.

OpenSSF Scorecard 7.4 / 10
7.4aggregate

Independent, tool-agnostic security assessment from the open-source OpenSSF Scorecard. Each check rewards a security practice, not a specific vendor's tool. Checks Scorecard could not determine are marked n/a and excluded from the security score (never counted as zero).Scorecard v5.5.0 · 2026-08-08 19:36 UTC

10Binary-Artifactsno binaries found in the repo
0Branch-Protectionbranch protection not enabled on development/release branches
10CI-Tests11 out of 11 merged PRs checked by a CI test -- score normalized to 10
2CII-Best-Practicesbadge detected: InProgress
5Code-ReviewFound 3/6 approved changesets -- score normalized to 5
10Contributorsproject has 21 contributing companies or organizations
10Dangerous-Workflowno dangerous workflow patterns detected
10Dependency-Update-Toolupdate tool detected
10Fuzzingproject is fuzzed
10Licenselicense file detected
10Maintained30 commit(s) and 7 issue activity found in the last 90 days -- score normalized to 10
10Packagingpackaging workflow detected
0Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
10SASTSAST tool is run on all commits
10Security-Policysecurity policy file detected
n/aSigned-Releasesno releases found
0Token-Permissionsdetected GitHub workflow tokens with excessive permissions
10Vulnerabilities0 existing vulnerabilities detected
Direct dependencies 2
RegistryPackageVersion constraintManifest
PyPIimportlib_resourcespyproject.toml
PyPItzdatapyproject.toml
All dependencies 67

Full resolved dependency set from the GitHub dependency graph: 0 direct and 67 indirect (transitive) packages. The transitive closure is complete when the repository commits a lockfile.

RegistryPackageVersionRelation
PyPIattacut1.0.6indirect
PyPIattaparse1.0.0indirect
PyPIbanditindirect
PyPIblackindirect
PyPIbpembindirect
PyPIbudoux0.7.0indirect
PyPIbuildindirect
PyPIbump-my-versionindirect
PyPIconlluindirect
PyPIcoverageindirect
PyPIdillindirect
PyPIemojiindirect
PyPIepitran1.26.0indirect
PyPIesuparindirect
PyPIfairseq-fixedindirect
PyPIfastaiindirect
PyPIfastcoref2.1.6indirect
PyPIflake8indirect
PyPIflake8-type-checkingindirect
PyPIfutureindirect
PyPIgensimindirect
PyPIhuggingface-hubindirect
PyPIkhamyoindirect
PyPIkhanaaindirect
PyPIlangdetectindirect
PyPImarisa-trieindirect
PyPImultielindirect
PyPImypyindirect
PyPInlpo3indirect
PyPInltkindirect
PyPInumpyindirect
PyPIonnxruntimeindirect
PyPIoskutindirect
PyPIpandasindirect
PyPIpanphon0.22.2indirect
PyPIphunspell0.1.6indirect
PyPIpyicuindirect
PyPIpyicu1.9.3indirect
PyPIpylintindirect
PyPIpython-crfsuite0.9.12indirect
PyPIpytzindirect
PyPIpyyamlindirect
PyPIrequestsindirect
PyPIruffindirect
PyPIsacremoses0.1.1indirect
PyPIsefr-cutindirect
PyPIsentence-transformersindirect
PyPIsentencepiece0.2.2indirect
PyPIsixindirect
PyPIspacyindirect
PyPIspacy-thai0.7.8indirect
PyPIsphinxindirect
PyPIsphinx-copybuttonindirect
PyPIsphinx-rtd-themeindirect
PyPIssg0.0.8indirect
PyPIsymspellpy6.10.0indirect
PyPIthai-nner0.3indirect
PyPItinydbindirect
PyPItltkindirect
PyPItorchindirect
PyPItoxindirect
PyPItqdmindirect
PyPItransformers5.14.1indirect
PyPIufal-chu-liu-edmonds1.0.3indirect
PyPIword2wordindirect
PyPIwtpsplit1.3.0indirect
PyPIwunsen0.0.3indirect
Dependency advisories 0

This repository publishes no package the index resolves, so its own dependency graph was assessed — 19 packages, which also include development and test pins that never ship: 0 carry known advisories, of which 0 are direct. 48 could not be assessed — no resolved version, an unsupported ecosystem, or beyond the reported package list.

No known advisories affect the assessed dependencies.

An advisory means the version recorded in the dependency graph falls inside an advisory’s affected range. Reachability is not analysed, and the graph includes development and test pins — a finding may concern tooling rather than shipped software.

Raw JSON report machine-readable

Scores are signals, not warranties. They reflect publicly visible practices on GitHub — not a code audit, and not a security guarantee.

Missing data is excluded and weights renormalized, never scored as zero. Methodology is versioned and open: metrics v2.5.0, schema v0.31.0 — full methodology · metrics wiki.

How one result sits in the wider record: aggregate statisticsPyPI.