Public record
Software health reportschema 0.31.0 · metrics 2.10.0 · 2026-08-05 05:44 UTC

piskvorky / gensim

Topic Modelling for Humans

PythonLGPL-2.1★ 16,479 stars⑂ 4,406 forkssince Feb 2011View on GitHub ↗
KindCommand-line toolhow this is determined

piskvorky/gensim holds a health index of 32 out of 100, placing it in the At Risk band. It scores highest on Community & Adoption (90/100) and lowest on Security (36/100). It was last updated 276 days ago. 3 contributors account for most of its recent work.

32
overall / 100
At Risk

Software health index

Metrics are grouped into weighted categories on one standardized 1–100 scale. Overall starts as their weighted mean, calibrated against the distribution of the public record so bands carry percentile meaning; when public evidence triggers the High-Risk Jurisdiction Policy, the rating is adjusted and receives an At Risk ceiling of 34.

32
Exceptional93-100The record's top tier (≈ top 5%); essentially all checked criteria met
Excellent80-92Strong across the board; minor gaps
Good65-79Healthy; gaps are limited and manageable
Moderate50-64Acceptable with notable gaps; review recommended
Weak35-49Material weaknesses across several areas
At Risk20-34Significant weaknesses; adoption warrants caution
Critical1-19Severe problems (abandoned, single-maintainer, no hygiene)
VitalityCommunity &AdoptionSustainability &GovernanceEngineeringQualitySecurityAI Readiness

Score profile

Each axis is a category. The shape matters more than the average — a healthy subject fills the whole shape, while a spike-and-crater profile means strength in one dimension is masking risk in another.

The weighted overall 60 is calibrated to 65 on the published index scale (record calibration 2026-08-02). High-Risk Jurisdiction Policy applies a 50% multiplier to weighted overall health and gives it an At Risk ceiling of 34.

Ownership

Radim ŘehůřekPersonal account
993 followers60 public repossince Feb 2011@pii-tools

This repository is owned by a personal account. A single-owner project carries more continuity risk than an organization-backed one.

Package ecosystems

RegistryPackageVersionDownloads / moVersionsLast publishTags
PyPIgensimpoints to another repo — not scored4.4.05,147,00180291 days agosingular-value-decompositionsvdlatent-semantic-indexinglsalsilatent-dirichlet-allocationldahierarchical-dirichlet-processhdprandom-projectionstfidfword2vec

Metrics by category

Vitality

Is the project alive — is code being written and are releases shipping?

38Weak · 21% of overall
How it's scored
3.6/36Push recencylast push 276 days ago
2.8/36Commit cadence4/52 weeks with commits
11.1/18Commit volume16 commits in the last year
5/10OpenSSF Scorecard: Maintained0 commit(s) and 6 issue activity found in the last 90 days -- score normalized to 5
Inputs used
commits_last_year16
human_commit_share0.91
days_since_last_push276
active_weeks_last_year4
How it's scored
27/27Ships releases44 releases published
16.2/36Release recencylatest release 292 days ago
12.6/27Release cadencea release every ~185.5 days
0/10OpenSSF Scorecard: Signed-Releasesno data
Inputs used
releases_count44
latest_release_tag4.4.0
releases_from_tagsno
days_since_latest_release292
mean_days_between_releases185.5
Excluded from scoring (no data or not applicable): OpenSSF Scorecard: Signed-Releases. Remaining weights renormalized.

Community & Adoption

Does the project have users, downloads, attention, and a welcoming setup for contributors?

90Excellent · 17% of overall

Popularity & adoption

100Exceptional
How it's scored
60/60Stars16,479 stars
25/25Forks4,406 forks
14.5/15Watchers406 watchers
Inputs used
forks4,406
stars16,479
watchers406
growth_stateunverified
growth_factor_pct100
growth_unverified_reasonno_history
How it's scored
22.5/22.5README
22.5/22.5Licenserecognized license (LGPL-2.1)
18/18CONTRIBUTING guide
0/13.5Code of conduct
7.2/7.2Issue template
0/6.3PR template
Inputs used
has_readmeyes
has_licenseyes
readme_badges5
has_contributingyes
has_issue_templateyes
has_code_of_conductno
readme_badge_servicesgithub.com, shields.io
has_pull_request_templateno

Sustainability & Governance

Will the project survive its people — bus factor, responsiveness, who backs it, and package upkeep?

68Good · 23% of overall
How it's scored
36/54Bus factor3 contributor(s) cover half of all commits
13.5/22.5Commit distributiontop contributor authored 40% of commits
13.5/13.5Contributor breadth99 contributors
10/10OpenSSF Scorecard: Contributorsproject has 46 contributing companies or organizations
Inputs used
bus_factor3
contributors_sampled99
top_contributor_share0.399
How it's scored
33.2/42Issue resolution79% of issues closed
21.6/30PR acceptance1,223/1,701 decided PRs merged
0/13Newcomer PR acceptance0/1 first-time contributors' PRs merged in 30d
3/15OpenSSF Scorecard: Code-ReviewFound 2/10 approved changesets -- score normalized to 2
Inputs used
merged_prs1,223
open_issues393
closed_issues1,488
prs_merged_7d0
prs_decided_7d0
prs_merged_30d0
prs_decided_30d1
issue_closed_ratio0.791
closed_unmerged_prs478
first_time_authors_30d1
first_time_prs_merged_30d0
first_time_prs_decided_30d1
How it's scored
10/30Ownership backingpersonal (user) account
0/20Verified domainnot applicable to user accounts
21.5/25Owner reach993 followers of piskvorky
25/25Track record60 public repos, account ~15 yr old
Inputs used
followers993
owner_typeUser
is_verified
owner_loginpiskvorky
public_repos60
account_age_days5,654
Excluded from scoring (no data or not applicable): Verified domain. Remaining weights renormalized.

Engineering Quality

Are baseline engineering and documentation practices in place?

72Good · 19% of overall
How it's scored
24/24CI workflows4 workflow(s)
24/24Tests present
0/16Linter config
0/9.6Pre-commit hooks
0/6.4.editorconfig
6/20OpenSSF Scorecard: CI-Tests1 out of 3 merged PRs checked by a CI test -- score normalized to 3
Inputs used
has_ciyes
has_testsyes
has_editorconfigno
has_linter_configno
has_precommit_configno

Documentation

100Exceptional
How it's scored
30/30README
25/25Documentation directory
15/15Documentation / homepage sitehttps://radimrehurek.com/gensim
10/10Repository description
10/10Topics15 topics
10/10Wiki
Inputs used
topicsgensim, topic-modeling, information-retrieval, machine-learning, natural-language-processing, nlp, data-science, python, data-mining, word2vec, word-embeddings, neural-network, document-similarity, word-similarity, fasttext
has_wikiyes
homepagehttps://radimrehurek.com/gensim
docs_sitehttps://radimrehurek.com/gensim
has_readmeyes
has_docs_diryes
has_descriptionyes

Security

Are visible security and supply-chain practices strong, without unresolved high-risk jurisdiction exposure?

36Weak · 16% of overall
How it's scored
7.5/7.5Binary-Artifactsno binaries found in the repo
0/7.5Branch-Protectionbranch protection not enabled on development/release branches
0.8/2.5CI-Tests1 out of 3 merged PRs checked by a CI test -- score normalized to 3
0/2.5CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
1.5/7.5Code-ReviewFound 2/10 approved changesets -- score normalized to 2
2.5/2.5Contributorsproject has 46 contributing companies or organizations
10/10Dangerous-Workflowno dangerous workflow patterns detected
7.5/7.5Dependency-Update-Toolupdate tool detected
0/5Fuzzingproject is not fuzzed
2.5/2.5Licenselicense file detected
3.8/7.5Maintained0 commit(s) and 6 issue activity found in the last 90 days -- score normalized to 5
0/5Packagingno data
0/5Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
0/5SASTSAST tool is not run on all commits -- score normalized to 0
2/5Security-Policysecurity policy file detected
0/7.5Signed-Releasesno data
0/7.5Token-Permissionsdetected GitHub workflow tokens with excessive permissions
0/7.5Vulnerabilities21 existing vulnerabilities detected
Inputs used
sourceopenssf_scorecard
checks_evaluated16
scorecard_versionv5.5.0
checks_inconclusive2
scorecard_aggregate4.1
high_risk_jurisdiction_cap34
high_risk_jurisdiction_multiplier50
security_posture_after_multiplier20
security_posture_before_jurisdiction41
Excluded from scoring (no data or not applicable): Packaging, Signed-Releases. Remaining weights renormalized. High-Risk Jurisdiction Policy applies a 50% multiplier and gives Security posture an At risk ceiling of 34.

Dependency advisories

100Exceptional
How it's scored
35/35Direct dependencies free of known advisoriesno direct dependency carries a known advisory
25/25Indirect dependencies free of known advisoriesno indirect dependency carries a known advisory
0/40No advisories left outstandingno advisory carries a publication date
Inputs used
sourceosv
advisories0
affected_packages0
assessed_packages4
unassessed_packages0
affected_by_severitynone
direct_affected_packages0
Excluded from scoring (no data or not applicable): No advisories left outstanding. Remaining weights renormalized. Matched the pypi:gensim@4.4.0 runtime dependency closure — what installing the published package pulls in — 4 packages. Reachability is not analyzed.

AI Readiness

How well is the repo equipped to be developed and maintained with AI coding agents? Carries a deliberately small weight (4%): agent tooling is a real maintenance signal, but a repository with none can still reach 100/100.

45Weak · 4% of overall
How it's scored
0/45Agent instructionsno CLAUDE.md / AGENTS.md / editor rules
0/15Machine-readable docs (llms.txt)
8.8/40Legible commit history15 of 91 human commits state their intent (structured subject or explanatory body)
Inputs used
has_llms_txtno
llms_txt_url
legible_history_share0.165
agent_instruction_files
agent_instruction_max_bytes
How it's scored
18/18One-command bootstrapdocs/src/Makefile
22/22Automated tests
0/11Lint / format config
0/11Static type checking
0/10Reproducible environment
0/10Demonstrated agent practiceno agent-authored commits among the last 100
8/8Automated maintenance9 of the last 100 commits are automated dependency updates
0/10OpenSSF Scorecard: Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
Inputs used
has_nixno
has_testsyes
lockfiles
has_dockerfileno
typed_languageno
bootstrap_filesdocs/src/Makefile
has_devcontainerno
has_linter_configno
typecheck_configs
agent_commit_share0
toolchain_manifests
dependency_bot_commit_share0.09
How it's scored
0/45Type-checkable codePython without a type-check config
51.5/55Manageable file sizes13/206 source files over 60KB
Inputs used
primary_languagePython
largest_source_bytes276,277
source_files_sampled206
oversized_source_files13
How it's scored
0/40API schema (OpenAPI/GraphQL/proto)not applicable to this kind of software
0/20MCP servernot applicable to this kind of software
40/40Runnable examplesexamples, notebooks
Inputs used
example_dirsexamples, notebooks
has_mcp_signalno
api_schema_files
interfaces_expected_of
Excluded from scoring (no data or not applicable): API schema (OpenAPI/GraphQL/proto), MCP server. Remaining weights renormalized.

Key facts

16,479GitHub stars
99contributors
16commits, last 12 months
276days since last push
44releases
3bus factor
393open issues
PyPIpackage ecosystems

Data collection warnings

  • Star history unavailable: GitHub GraphQL error: Resource not accessible by personal access token
  • pypi package 'gensim' points at a different repository (https://github.com/RaRe-Technologies/gensim); excluded from ecosystem scoring

More detail

Star and fork history 0 ★ / 4,406 ⇿
0Stars
4,406Forks
11Releases

When each star and fork was added, collected from GitHub and bucketed by day. Cumulative growth sits directly above the daily additions it is made of, so the two read against each other: steady organic accretion looks nothing like an abrupt, short-lived burst. Where that difference is measurable, it is reported as growth authenticity.

Only the most recent history is shown — this repository exceeds the collection window, so the earliest history is not captured.

3,0003,5004,0004,5004,406142020-082023-082026-08
Major 1Minor 4Patch 4

Each point covers 6 days.

OpenSSF Scorecard 4.1 / 10
4.1aggregate

Independent, tool-agnostic security assessment from the open-source OpenSSF Scorecard. Each check rewards a security practice, not a specific vendor's tool. Checks Scorecard could not determine are marked n/a and excluded from the security score (never counted as zero).Scorecard v5.5.0 · 2026-08-05 05:44 UTC

10Binary-Artifactsno binaries found in the repo
0Branch-Protectionbranch protection not enabled on development/release branches
3CI-Tests1 out of 3 merged PRs checked by a CI test -- score normalized to 3
0CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
2Code-ReviewFound 2/10 approved changesets -- score normalized to 2
10Contributorsproject has 46 contributing companies or organizations
10Dangerous-Workflowno dangerous workflow patterns detected
10Dependency-Update-Toolupdate tool detected
0Fuzzingproject is not fuzzed
10Licenselicense file detected
5Maintained0 commit(s) and 6 issue activity found in the last 90 days -- score normalized to 5
n/aPackagingpackaging workflow not detected
0Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
0SASTSAST tool is not run on all commits -- score normalized to 0
4Security-Policysecurity policy file detected
n/aSigned-Releasesno releases found
0Token-Permissionsdetected GitHub workflow tokens with excessive permissions
0Vulnerabilities21 existing vulnerabilities detected
All dependencies 16

Full resolved dependency set from the GitHub dependency graph: 0 direct and 16 indirect (transitive) packages. The transitive closure is complete when the repository commits a lockfile.

RegistryPackageVersionRelation
PyPIannoy1.16.2indirect
PyPImemory-profiler0.55.0indirect
PyPInltk3.4.5indirect
PyPInmslib2.1.1indirect
PyPIpandas1.2.3indirect
PyPIpot0.8.1indirect
PyPIpyro44.77indirect
PyPIscikit-learn0.24.1indirect
PyPIscipyindirect
PyPIsmart-openindirect
PyPIsphinx3.5.2indirect
PyPIsphinx-gallery0.8.2indirect
PyPIsphinxcontrib-napoleon0.7indirect
PyPIsphinxcontrib-programoutput0.15indirect
PyPIstatsmodels0.12.2indirect
PyPItestfixtures6.17.1indirect
Dependency advisories 0

Installing pypi:gensim@4.4.0 pulls in 4 packages, direct and transitive: 0 carry known advisories, of which 0 are direct dependencies.

No known advisories affect the assessed dependencies.

An advisory means the version recorded in the dependency graph falls inside an advisory’s affected range. Reachability is not analysed, and the graph includes development and test pins — a finding may concern tooling rather than shipped software.

Raw JSON report machine-readable

Feedback

Spotted something off in this report, or have thoughts to share? Wrong measurements, missed tooling, ideas, questions — anything is welcome. Every message is read and gets a response.

The message is kept through sign-in.

Scores are signals, not warranties. They reflect publicly visible practices on GitHub — not a code audit, and not a security guarantee.

Missing data is excluded and weights renormalized, never scored as zero. Methodology is versioned and open: metrics v2.10.0, schema v0.31.0 — full methodology · metrics wiki.

How one result sits in the wider record: aggregate statisticsPyPI.