# david-smejkal/wiki2txt — health index 28/100 (At Risk)

> Inspection of the public repository david-smejkal/wiki2txt by inspect.software. It holds a health index of 28 out of 100, placing it in the At Risk band. The index is a signal derived from publicly visible practice, not a warranty, a security audit, or an endorsement, and it is independent of payment.

- Repository: https://github.com/david-smejkal/wiki2txt
- Report page: https://inspect.software/software/david-smejkal/wiki2txt
- Full JSON report: https://inspect.software/api/repositories/david-smejkal/wiki2txt/report
- Badge: https://raw.githubusercontent.com/inspect-software/badges/main/v1/d/david-smejkal/wiki2txt.svg
- Inspected: 2026-08-13
- Methodology: metrics 2.10.0, report schema 0.31.0 (https://inspect.software/methodology.md)

## Category summary

| Category | Weight | Index | Band |
|---|---|---|---|
| Vitality | 21% | 18/100 | Critical |
| Community & Adoption | 17% | 30/100 | At Risk |
| Sustainability & Governance | 23% | 47/100 | Weak |
| Engineering Quality | 19% | 54/100 | Moderate |
| Security | 16% | 21/100 | At Risk |
| AI Readiness | 4% | 27/100 | At Risk |

The weighted overall 34 is calibrated to 28 on the published index scale (record calibration 2026-08-02).

## Repository facts

| Field | Value |
|---|---|
| Description | A tool to extract plain (unformatted) multilingual / language-agnostic text, redirects, links and categories from wikipedia backups (dumps). Designed to prepare clean training data for AI Training / Machine Learning software. |
| Primary language | Python |
| License | GPL-2.0 |
| Stars | 7 |
| Forks | 1 |
| Created | 2021-12-01 |
| Last push | 2025-03-31 |
| Latest release | v0.7.0 |
| Commits (last year) | 0 |
| Bus factor | 1 |
| Topics | wiki-to-txt, wiki-to-text, wiki-to-plaintext, wikidump-to-txt, wikidump-to-plaintext, wiki-parser, wikidump-parser, ai-learning-tool, tool-for-ai, wikidumps-parser, wiki2plaintext, ai-learning, data-parser-for-ai, data-for-robots, plaintext-data-for-ai, wikipedia-to-txt, machine-learning-tool, machine-learning, training-data, ai-training |

## Vitality — 18/100 (Critical)

Is the project alive — is code being written and are releases shipping? Weight: 21% of the overall index.

### Development activity — 1/100 (Critical)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Push recency | not met | 0 / 36 | last push 500 days ago |
| Commit cadence | not met | 0 / 36 | 0/52 weeks with commits |
| Commit volume | not met | 0 / 18 | 0 commits in the last year |
| OpenSSF Scorecard: Maintained | excluded | 0 / 10 | OpenSSF Scorecard unavailable |

### Release discipline — 44/100 (Weak)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Ships releases | met | 27 / 27 | 3 releases published |
| Release recency | partial | 7.2 / 36 | latest release 503 days ago |
| Release cadence | partial | 5.4 / 27 | a release every ~604 days |
| OpenSSF Scorecard: Signed-Releases | excluded | 0 / 10 | OpenSSF Scorecard unavailable |

### Abandonment — 100/100 (Exceptional)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Project is still maintained | met | 100 / 100 | no human commit for 509 days, with nothing left unanswered; held at dormant by nothing open to answer |

## Community & Adoption — 30/100 (At Risk)

Does the project have users, downloads, attention, and a welcoming setup for contributors? Weight: 17% of the overall index.

### Popularity & adoption — 13/100 (Critical)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Stars | partial | 12.6 / 60 | 7 stars |
| Forks | not met | 0 / 25 | 1 forks |
| Watchers | not met | 0 / 15 | 2 watchers |

### Community health — 50/100 (Moderate)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| README | met | 22.5 / 22.5 | — |
| License | met | 22.5 / 22.5 | recognized license (GPL-2.0) |
| CONTRIBUTING guide | not met | 0 / 18 | — |
| Code of conduct | not met | 0 / 13.5 | — |
| Issue template | not met | 0 / 7.2 | — |
| PR template | not met | 0 / 6.3 | — |

## Sustainability & Governance — 47/100 (Weak)

Will the project survive its people — bus factor, responsiveness, who backs it, and package upkeep? Weight: 23% of the overall index.

### Maintainer resilience (bus factor) — 12/100 (Critical)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Bus factor | partial | 9 / 54 | 1 contributor(s) cover half of all commits |
| Commit distribution | not met | 0 / 22.5 | top contributor authored 100% of commits |
| Contributor breadth | partial | 1.4 / 13.5 | 1 contributors |
| OpenSSF Scorecard: Contributors | excluded | 0 / 10 | OpenSSF Scorecard unavailable |

### Issue & PR responsiveness — 100/100 (Exceptional)

Excluded from scoring (no data or not applicable): PR acceptance, Newcomer PR acceptance. Remaining weights renormalized.

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Issue resolution | met | 42 / 42 | 100% of issues closed |
| PR acceptance | excluded | 0 / 30 | no decided pull requests or no data |
| Newcomer PR acceptance | excluded | 0 / 13 | no first-time contributor's PR decided in 30d |
| OpenSSF Scorecard: Code-Review | excluded | 0 / 15 | OpenSSF Scorecard unavailable |

### Ownership & stewardship — 36/100 (Weak)

Excluded from scoring (no data or not applicable): Verified domain. Remaining weights renormalized.

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Ownership backing | partial | 10 / 30 | personal (user) account |
| Verified domain | excluded | 0 / 20 | not applicable to user accounts |
| Owner reach | partial | 2.2 / 25 | 1 followers of david-smejkal |
| Track record | partial | 16.4 / 25 | 8 public repos, account ~4 yr old |

## Engineering Quality — 54/100 (Moderate)

Are baseline engineering and documentation practices in place? Weight: 19% of the overall index.

### Engineering practices — 50/100 (Moderate)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| CI workflows | not met | 0 / 24 | — |
| Tests present | met | 24 / 24 | — |
| Linter config | met | 16 / 16 | tox.ini |
| Pre-commit hooks | not met | 0 / 9.6 | — |
| .editorconfig | not met | 0 / 6.4 | — |
| OpenSSF Scorecard: CI-Tests | excluded | 0 / 20 | OpenSSF Scorecard unavailable |

### Documentation — 60/100 (Moderate)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| README | met | 30 / 30 | — |
| Documentation directory | not met | 0 / 25 | — |
| Documentation / homepage site | not met | 0 / 15 | — |
| Repository description | met | 10 / 10 | — |
| Topics | met | 10 / 10 | 20 topics |
| Wiki | met | 10 / 10 | — |

## Security — 21/100 (At Risk)

Are visible security and supply-chain practices strong, with no malicious dependency and no unresolved high-risk jurisdiction exposure? Weight: 16% of the overall index.

### Security posture — 1/100 (Critical)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Security policy (SECURITY.md) | not met | 0 / 30 | — |
| Dependabot config | not met | 0 / 25 | — |
| Dependency lockfiles | not met | 0 / 25 | — |
| CodeQL workflow | not met | 0 / 20 | — |

### Dependency advisories — 100/100 (Exceptional)

Excluded from scoring (no data or not applicable): Indirect dependencies free of known advisories, No advisories left outstanding. Remaining weights renormalized. Matched 15 resolved dependencies against OSV; 4 could not be assessed (no resolved version, an unsupported ecosystem, or beyond the reported package list). This repository publishes no package the index resolves, so the repository dependency graph was assessed instead. That graph mixes development and test pins with shipped dependencies, so only the declared runtime dependencies are scored; transitive findings are reported as context and excluded from the score. Reachability is not analyzed.

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Direct dependencies free of known advisories | met | 35 / 35 | no direct dependency carries a known advisory |
| Indirect dependencies free of known advisories | excluded | 0 / 25 | transitive set not separable from development and test dependencies in this scope |
| No advisories left outstanding | excluded | 0 / 40 | no advisory carries a publication date |

### Malicious dependencies — 100/100 (Exceptional)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| No dependency reported as a malicious package | met | 100 / 100 | no dependency is reported as a malicious package |

### High-Risk Jurisdiction Exposure — 100/100 (Exceptional)

Only high-confidence self-published location evidence affects this multiplier. Ambiguous matches are review-only; country evidence is not proof of nationality, citizenship, legal registration, malicious intent, or sanctions status.

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Policy exposure multiplier | met | 100 / 100 | no confirmed policy-scope location match |

## AI Readiness — 27/100 (At Risk)

How well is the repo equipped to be developed and maintained with AI coding agents? Carries a deliberately small weight: agent tooling is a real maintenance signal, but its absence must never gate the top of the scale (calibration saturates at raw 91, so 100/100 remains reachable with AI Readiness at zero).

### Agent context & guidance — 1/100 (Critical)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Agent instructions | not met | 0 / 45 | no CLAUDE.md / AGENTS.md / editor rules |
| Machine-readable docs (llms.txt) | not met | 0 / 15 | — |
| Legible commit history | partial | 1.1 / 40 | 2 of 100 human commits state their intent (structured subject or explanatory body) |

### Verify loop (build / test / typecheck) — 37/100 (Weak)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| One-command bootstrap | not met | 0 / 18 | — |
| Automated tests | met | 22 / 22 | — |
| Lint / format config | met | 11 / 11 | tox.ini |
| Static type checking | not met | 0 / 11 | — |
| Reproducible environment | not met | 0 / 10 | — |
| Demonstrated agent practice | not met | 0 / 10 | no agent-authored commits among the last 100 |
| Automated maintenance | not met | 0 / 8 | no automated dependency updates observed |
| OpenSSF Scorecard: Pinned-Dependencies | excluded | 0 / 10 | OpenSSF Scorecard unavailable |

### Code legibility for models — 55/100 (Moderate)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Type-checkable code | not met | 0 / 45 | Python without a type-check config |
| Manageable file sizes | met | 55 / 55 | 0/9 source files over 60KB |

## Collection warnings

- Star history unavailable: GitHub GraphQL error: Resource not accessible by personal access token
- OpenSSF Scorecard timed out after 240s; skipping Scorecard checks

## What this report is

inspect.software measures public repositories against a single published, versioned methodology and reports the result as a 1–100 health index. Findings are signals derived from what is visible in public repository data — they are not a security audit, a warranty, or an endorsement, and no result depends on payment.

- Methodology, formulas and weights: https://inspect.software/methodology.md
- What a result does and does not claim: https://inspect.software/wiki/signals-not-warranties.md
- Rating bands: https://inspect.software/wiki/scoring-bands.md
- Corrections: https://inspect.software/contact.md
- Index of all documents: https://inspect.software/llms.txt
