Public record
Software health reportschema 0.31.0 · metrics 2.10.0 · 2026-08-13 15:27 UTC

david-smejkal / wiki2txt

A tool to extract plain (unformatted) multilingual / language-agnostic text, redirects, links and categories from wikipedia backups (dumps). Designed to prepare clean training data for AI Training / Machine Learning software.

PythonGPL-2.0★ 7 stars⑂ 1 forksince Dec 2021View on GitHub ↗

david-smejkal/wiki2txt holds a health index of 28 out of 100, placing it in the At Risk band. It scores highest on Engineering Quality (54/100) and lowest on Vitality (18/100). It was last updated 500 days ago. A single contributor accounts for most of its recent work.

28
overall / 100
At Risk

Software health index

Metrics are grouped into weighted categories on one standardized 1–100 scale. Overall starts as their weighted mean, calibrated against the distribution of the public record so bands carry percentile meaning; when public evidence triggers the High-Risk Jurisdiction Policy, the rating is adjusted and receives an At Risk ceiling of 34.

28
Exceptional93-100The record's top tier (≈ top 5%); essentially all checked criteria met
Excellent80-92Strong across the board; minor gaps
Good65-79Healthy; gaps are limited and manageable
Moderate50-64Acceptable with notable gaps; review recommended
Weak35-49Material weaknesses across several areas
At Risk20-34Significant weaknesses; adoption warrants caution
Critical1-19Severe problems (abandoned, single-maintainer, no hygiene)
VitalityCommunity &AdoptionSustainability &GovernanceEngineeringQualitySecurityAI Readiness

Score profile

Each axis is a category. The shape matters more than the average — a healthy subject fills the whole shape, while a spike-and-crater profile means strength in one dimension is masking risk in another.

The weighted overall 34 is calibrated to 28 on the published index scale (record calibration 2026-08-02).

Ownership

David SmejkalPersonal account
1 follower8 public repossince Nov 2021

This repository is owned by a personal account. A single-owner project carries more continuity risk than an organization-backed one.

Metrics by category

Vitality

Is the project alive — is code being written and are releases shipping?

18Critical · 21% of overall
How it's scored
0/36Push recencylast push 500 days ago
0/36Commit cadence0/52 weeks with commits
0/18Commit volume0 commits in the last year
0/10OpenSSF Scorecard: Maintainedno data
Inputs used
commits_last_year0
human_commit_share1
days_since_last_push500
active_weeks_last_year0
How it's scored
27/27Ships releases3 releases published
7.2/36Release recencylatest release 503 days ago
5.4/27Release cadencea release every ~604 days
0/10OpenSSF Scorecard: Signed-Releasesno data
Inputs used
releases_count3
latest_release_tagv0.7.0
releases_from_tagsno
days_since_latest_release503
mean_days_between_releases604

Community & Adoption

Does the project have users, downloads, attention, and a welcoming setup for contributors?

30At Risk · 17% of overall
How it's scored
12.6/60Stars7 stars
0/25Forks1 forks
0/15Watchers2 watchers
Inputs used
forks1
stars7
watchers2
growth_stateunverified
growth_factor_pct100
growth_unverified_reasonno_history
How it's scored
22.5/22.5README
22.5/22.5Licenserecognized license (GPL-2.0)
0/18CONTRIBUTING guide
0/13.5Code of conduct
0/7.2Issue template
0/6.3PR template
Inputs used
has_readmeyes
has_licenseyes
readme_badges0
has_contributingno
has_issue_templateno
has_code_of_conductno
readme_badge_services
has_pull_request_templateno

Sustainability & Governance

Will the project survive its people — bus factor, responsiveness, who backs it, and package upkeep?

47Weak · 23% of overall
How it's scored
9/54Bus factor1 contributor(s) cover half of all commits
0/22.5Commit distributiontop contributor authored 100% of commits
1.4/13.5Contributor breadth1 contributors
0/10OpenSSF Scorecard: Contributorsno data
Inputs used
bus_factor1
contributors_sampled1
top_contributor_share1
How it's scored
42/42Issue resolution100% of issues closed
0/30PR acceptanceno decided pull requests or no data
0/13Newcomer PR acceptanceno first-time contributor's PR decided in 30d
0/15OpenSSF Scorecard: Code-Reviewno data
Inputs used
merged_prs0
open_issues0
closed_issues1
prs_merged_7d0
prs_decided_7d0
prs_merged_30d0
prs_decided_30d0
issue_closed_ratio1
closed_unmerged_prs0
first_time_authors_30d0
first_time_prs_merged_30d0
first_time_prs_decided_30d0
Excluded from scoring (no data or not applicable): PR acceptance, Newcomer PR acceptance. Remaining weights renormalized.
How it's scored
10/30Ownership backingpersonal (user) account
0/20Verified domainnot applicable to user accounts
2.2/25Owner reach1 followers of david-smejkal
16.4/25Track record8 public repos, account ~4 yr old
Inputs used
followers1
owner_typeUser
is_verified
owner_logindavid-smejkal
public_repos8
account_age_days1,718
Excluded from scoring (no data or not applicable): Verified domain. Remaining weights renormalized.

Engineering Quality

Are baseline engineering and documentation practices in place?

54Moderate · 19% of overall
How it's scored
0/24CI workflows
24/24Tests present
16/16Linter configtox.ini
0/9.6Pre-commit hooks
0/6.4.editorconfig
0/20OpenSSF Scorecard: CI-Testsno data
Inputs used
has_cino
has_testsyes
has_editorconfigno
has_linter_configyes
has_precommit_configno

Documentation

60Moderate
How it's scored
30/30README
0/25Documentation directory
0/15Documentation / homepage site
10/10Repository description
10/10Topics20 topics
10/10Wiki
Inputs used
topicswiki-to-txt, wiki-to-text, wiki-to-plaintext, wikidump-to-txt, wikidump-to-plaintext, wiki-parser, wikidump-parser, ai-learning-tool, tool-for-ai, wikidumps-parser, wiki2plaintext, ai-learning, data-parser-for-ai, data-for-robots, plaintext-data-for-ai, wikipedia-to-txt, machine-learning-tool, machine-learning, training-data, ai-training
has_wikiyes
homepage
docs_site
has_readmeyes
has_docs_dirno
has_descriptionyes

Security

Are visible security and supply-chain practices strong, without unresolved high-risk jurisdiction exposure?

21At Risk · 16% of overall
How it's scored
0/30Security policy (SECURITY.md)
0/25Dependabot config
0/25Dependency lockfiles
0/20CodeQL workflow
Inputs used
sourcefile_signals
lockfiles
manifestsrequirements-frozen.txt, requirements-test.txt, requirements.txt
has_codeql_workflowno
has_security_policyno
has_dependabot_configno

Dependency advisories

100Exceptional
How it's scored
35/35Direct dependencies free of known advisoriesno direct dependency carries a known advisory
0/25Indirect dependencies free of known advisoriestransitive set not separable from development and test dependencies in this scope
0/40No advisories left outstandingno advisory carries a publication date
Inputs used
sourceosv
advisories10
affected_packages4
assessed_packages15
unassessed_packages4
affected_by_severityhigh 1, moderate 3
direct_affected_packages0
Excluded from scoring (no data or not applicable): Indirect dependencies free of known advisories, No advisories left outstanding. Remaining weights renormalized. Matched 15 resolved dependencies against OSV. 4 could not be assessed — no resolved version, an unsupported ecosystem, or beyond the reported package list. This repository publishes no package the index resolves, so the repository dependency graph was assessed instead. That graph mixes development and test pins with shipped dependencies, so only the declared runtime dependencies are scored; transitive findings are reported as context and excluded from the score. Reachability is not analyzed.

AI Readiness

How well is the repo equipped to be developed and maintained with AI coding agents? Carries a deliberately small weight (4%): agent tooling is a real maintenance signal, but a repository with none can still reach 100/100.

27At Risk · 4% of overall
How it's scored
0/45Agent instructionsno CLAUDE.md / AGENTS.md / editor rules
0/15Machine-readable docs (llms.txt)
1.1/40Legible commit history2 of 100 human commits state their intent (structured subject or explanatory body)
Inputs used
has_llms_txtno
llms_txt_url
legible_history_share0.02
agent_instruction_files
agent_instruction_max_bytes
How it's scored
0/18One-command bootstrap
22/22Automated tests
11/11Lint / format configtox.ini
0/11Static type checking
0/10Reproducible environment
0/10Demonstrated agent practiceno agent-authored commits among the last 100
0/8Automated maintenanceno automated dependency updates observed
0/10OpenSSF Scorecard: Pinned-Dependenciesno data
Inputs used
has_nixno
has_testsyes
lockfiles
has_dockerfileno
typed_languageno
bootstrap_files
has_devcontainerno
has_linter_configyes
typecheck_configs
agent_commit_share0
toolchain_manifests
dependency_bot_commit_share0
How it's scored
0/45Type-checkable codePython without a type-check config
55/55Manageable file sizes0/9 source files over 60KB
Inputs used
primary_languagePython
largest_source_bytes43,396
source_files_sampled9
oversized_source_files0

Key facts

7GitHub stars
1contributors
0commits, last 12 months
500days since last push
3releases
1bus factor
0open issues
PyPIpackage ecosystems

Data collection warnings

  • Star history unavailable: GitHub GraphQL error: Resource not accessible by personal access token
  • OpenSSF Scorecard timed out after 240s; skipping Scorecard checks

More detail

All dependencies 19

Full resolved dependency set from the GitHub dependency graph: 0 direct and 19 indirect (transitive) packages. The transitive closure is complete when the repository commits a lockfile.

RegistryPackageVersionRelation
PyPIcachetools5.5.2indirect
PyPIchardet5.2.0indirect
PyPIcolorama0.4.6indirect
PyPIdistlib0.3.9indirect
PyPIfilelock3.18.0indirect
PyPIiniconfig2.1.0indirect
PyPIlxmlindirect
PyPIlxml5.3.1indirect
PyPIpackaging24.2indirect
PyPIplatformdirs4.3.7indirect
PyPIpluggy1.5.0indirect
PyPIpyproject-api1.9.0indirect
PyPIpytestindirect
PyPIpytest8.3.5indirect
PyPIregexindirect
PyPIregex2024.11.6indirect
PyPItoxindirect
PyPItox4.25.0indirect
PyPIvirtualenv20.29.3indirect
Dependency advisories 4

This repository publishes no package the index resolves, so its own dependency graph was assessed — 15 packages, which also include development and test pins that never ship: 4 carry known advisories, of which 0 are direct. 4 could not be assessed — no resolved version, an unsupported ecosystem, or beyond the reported package list.

PackageVersionRelationSeverityAdvisoriesFixed in
lxml5.3.1indirecthigh26.1.0
filelock3.18.0indirectmoderate43.20.3
pytest8.3.5indirectmoderate29.0.3
virtualenv20.29.3indirectmoderate220.36.1

An advisory means the version recorded in the dependency graph falls inside an advisory’s affected range. Reachability is not analysed, and the graph includes development and test pins — a finding may concern tooling rather than shipped software.

Raw JSON report machine-readable

Feedback

Spotted something off in this report, or have thoughts to share? Wrong measurements, missed tooling, ideas, questions — anything is welcome. Every message is read and gets a response.

The message is kept through sign-in.

Scores are signals, not warranties. They reflect publicly visible practices on GitHub — not a code audit, and not a security guarantee.

Missing data is excluded and weights renormalized, never scored as zero. Methodology is versioned and open: metrics v2.10.0, schema v0.31.0 — full methodology · metrics wiki.

How one result sits in the wider record: aggregate statistics.