# JohnSnowLabs/spark-nlp — health index 81/100 (Excellent)

> Inspection of the public repository JohnSnowLabs/spark-nlp by inspect.software. It holds a health index of 81 out of 100, placing it in the Excellent band. The index is a signal derived from publicly visible practice, not a warranty, a security audit, or an endorsement, and it is independent of payment.

- Repository: https://github.com/JohnSnowLabs/spark-nlp
- Report page: https://inspect.software/software/JohnSnowLabs/spark-nlp
- Full JSON report: https://inspect.software/api/repositories/JohnSnowLabs/spark-nlp/report
- Badge: https://raw.githubusercontent.com/inspect-software/badges/main/v1/j/JohnSnowLabs/spark-nlp.svg
- Inspected: 2026-08-13
- Methodology: metrics 2.10.0, report schema 0.31.0 (https://inspect.software/methodology.md)

## Category summary

| Category | Weight | Index | Band |
|---|---|---|---|
| Vitality | 21% | 89/100 | Excellent |
| Community & Adoption | 17% | 84/100 | Excellent |
| Sustainability & Governance | 23% | 79/100 | Good |
| Engineering Quality | 19% | 54/100 | Moderate |
| Security | 16% | 40/100 | Weak |
| AI Readiness | 4% | 39/100 | Weak |

The weighted overall 69 is calibrated to 81 on the published index scale (record calibration 2026-08-02).

## Repository facts

| Field | Value |
|---|---|
| Description | State of the Art Natural Language Processing |
| Homepage | https://sparknlp.org/ |
| Primary language | Scala |
| License | Apache-2.0 |
| Stars | 4,154 |
| Forks | 743 |
| Created | 2017-09-24 |
| Last push | 2026-08-11 |
| Latest release | 6.4.2 |
| Commits (last year) | 170 |
| Bus factor | 2 |
| Topics | nlp, natural-language-processing, spark, pyspark, named-entity-recognition, sentiment-analysis, lemmatizer, spell-checker, entity-extraction, part-of-speech-tagger, bert, transformers, tensorflow, language-detection, machine-translation, text-classification, llm, question-answering, llamacpp, onnx |

## Vitality — 89/100 (Excellent)

Is the project alive — is code being written and are releases shipping? Weight: 21% of the overall index.

### Development activity — 82/100 (Excellent)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Push recency | met | 36 / 36 | last push 1 days ago |
| Commit cadence | partial | 20.1 / 36 | 29/52 weeks with commits |
| Commit volume | met | 18 / 18 | 170 commits in the last year |
| OpenSSF Scorecard: Maintained | excluded | 0 / 10 | OpenSSF Scorecard unavailable |

### Release discipline — 100/100 (Exceptional)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Ships releases | met | 27 / 27 | 100 releases published |
| Release recency | met | 36 / 36 | latest release 49 days ago |
| Release cadence | met | 27 / 27 | a release every ~27.2 days |
| OpenSSF Scorecard: Signed-Releases | excluded | 0 / 10 | OpenSSF Scorecard unavailable |

### Abandonment — 100/100 (Exceptional)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Project is still maintained | met | 100 / 100 | last human commit 58 days ago |

## Community & Adoption — 84/100 (Excellent)

Does the project have users, downloads, attention, and a welcoming setup for contributors? Weight: 17% of the overall index.

### Popularity & adoption — 94/100 (Exceptional)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Stars | partial | 58.7 / 60 | 4,154 stars |
| Forks | partial | 23.9 / 25 | 743 forks |
| Watchers | partial | 11 / 15 | 97 watchers |

### Community health — 72/100 (Good)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| README | met | 22.5 / 22.5 | — |
| License | met | 22.5 / 22.5 | recognized license (Apache-2.0) |
| CONTRIBUTING guide | not met | 0 / 18 | — |
| Code of conduct | met | 13.5 / 13.5 | — |
| Issue template | not met | 0 / 7.2 | — |
| PR template | met | 6.3 / 6.3 | — |

## Sustainability & Governance — 79/100 (Good)

Will the project survive its people — bus factor, responsiveness, who backs it, and package upkeep? Weight: 23% of the overall index.

### Maintainer resilience (bus factor) — 57/100 (Moderate)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Bus factor | partial | 25.2 / 54 | 2 contributor(s) cover half of all commits |
| Commit distribution | partial | 12.9 / 22.5 | top contributor authored 42% of commits |
| Contributor breadth | met | 13.5 / 13.5 | 92 contributors |
| OpenSSF Scorecard: Contributors | excluded | 0 / 10 | OpenSSF Scorecard unavailable |

### Issue & PR responsiveness — 95/100 (Exceptional)

Excluded from scoring (no data or not applicable): Newcomer PR acceptance. Remaining weights renormalized.

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Issue resolution | partial | 41.2 / 42 | 98% of issues closed |
| PR acceptance | partial | 27.3 / 30 | 12424/13674 decided PRs merged |
| Newcomer PR acceptance | excluded | 0 / 13 | no first-time contributor's PR decided in 30d |
| OpenSSF Scorecard: Code-Review | excluded | 0 / 15 | OpenSSF Scorecard unavailable |

### Ownership & stewardship — 88/100 (Excellent)

Excluded from scoring (no data or not applicable): Verified domain. Remaining weights renormalized.

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Ownership backing | met | 30 / 30 | organization-owned |
| Verified domain | excluded | 0 / 20 | verified-domain status not read for this organization |
| Owner reach | partial | 18.7 / 25 | 402 followers of JohnSnowLabs |
| Track record | partial | 21.9 / 25 | 22 public repos, account ~10 yr old |

## Engineering Quality — 54/100 (Moderate)

Are baseline engineering and documentation practices in place? Weight: 19% of the overall index.

### Engineering practices — 30/100 (At Risk)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| CI workflows | met | 24 / 24 | 4 workflow(s) |
| Tests present | not met | 0 / 24 | — |
| Linter config | not met | 0 / 16 | — |
| Pre-commit hooks | not met | 0 / 9.6 | — |
| .editorconfig | not met | 0 / 6.4 | — |
| OpenSSF Scorecard: CI-Tests | excluded | 0 / 20 | OpenSSF Scorecard unavailable |

### Documentation — 90/100 (Excellent)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| README | met | 30 / 30 | — |
| Documentation directory | met | 25 / 25 | — |
| Documentation / homepage site | met | 15 / 15 | https://sparknlp.org/ |
| Repository description | met | 10 / 10 | — |
| Topics | met | 10 / 10 | 20 topics |
| Wiki | not met | 0 / 10 | — |

## Security — 40/100 (Weak)

Are visible security and supply-chain practices strong, with no malicious dependency and no unresolved high-risk jurisdiction exposure? Weight: 16% of the overall index.

### Security posture — 25/100 (At Risk)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Security policy (SECURITY.md) | not met | 0 / 30 | — |
| Dependabot config | not met | 0 / 25 | — |
| Dependency lockfiles | met | 25 / 25 | Gemfile.lock, yarn.lock |
| CodeQL workflow | not met | 0 / 20 | — |

### Dependency advisories — 100/100 (Exceptional)

Excluded from scoring (no data or not applicable): Indirect dependencies free of known advisories, No advisories left outstanding. Remaining weights renormalized. Matched 796 resolved dependencies against OSV; 24 could not be assessed (no resolved version, an unsupported ecosystem, or beyond the reported package list). This repository publishes no package the index resolves, so the repository dependency graph was assessed instead. That graph mixes development and test pins with shipped dependencies, so only the declared runtime dependencies are scored; transitive findings are reported as context and excluded from the score. Reachability is not analyzed.

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Direct dependencies free of known advisories | met | 35 / 35 | no direct dependency carries a known advisory |
| Indirect dependencies free of known advisories | excluded | 0 / 25 | transitive set not separable from development and test dependencies in this scope |
| No advisories left outstanding | excluded | 0 / 40 | no advisory carries a publication date |

### Malicious dependencies — 100/100 (Exceptional)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| No dependency reported as a malicious package | met | 100 / 100 | no dependency is reported as a malicious package |

### High-Risk Jurisdiction Exposure — 100/100 (Exceptional)

Only high-confidence self-published location evidence affects this multiplier. Ambiguous matches are review-only; country evidence is not proof of nationality, citizenship, legal registration, malicious intent, or sanctions status.

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Policy exposure multiplier | met | 100 / 100 | no confirmed policy-scope location match |

## AI Readiness — 39/100 (Weak)

How well is the repo equipped to be developed and maintained with AI coding agents? Carries a deliberately small weight: agent tooling is a real maintenance signal, but its absence must never gate the top of the scale (calibration saturates at raw 91, so 100/100 remains reachable with AI Readiness at zero).

### Agent context & guidance — 17/100 (Critical)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Agent instructions | not met | 0 / 45 | no CLAUDE.md / AGENTS.md / editor rules |
| Machine-readable docs (llms.txt) | not met | 0 / 15 | — |
| Legible commit history | partial | 17.1 / 40 | 32 of 100 human commits state their intent (structured subject or explanatory body) |

### Verify loop (build / test / typecheck) — 32/100 (At Risk)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| One-command bootstrap | not met | 0 / 18 | — |
| Automated tests | not met | 0 / 22 | — |
| Lint / format config | not met | 0 / 11 | — |
| Static type checking | met | 11 / 11 | Scala (statically typed) |
| Reproducible environment | met | 10 / 10 | lockfile |
| Demonstrated agent practice | partial | 8 / 10 | 4 of the last 100 commits agent-authored or agent-credited |
| Automated maintenance | not met | 0 / 8 | no automated dependency updates observed |
| OpenSSF Scorecard: Pinned-Dependencies | excluded | 0 / 10 | OpenSSF Scorecard unavailable |

### Code legibility for models — 99/100 (Exceptional)

| Criterion | Status | Points | Detail |
|---|---|---|---|
| Type-checkable code | met | 45 / 45 | Scala (statically typed) |
| Manageable file sizes | partial | 54.2 / 55 | 1/73 source files over 60KB |

## Collection warnings

- Star history unavailable: GitHub GraphQL error: Resource not accessible by personal access token
- File tree truncated by GitHub API; file-based signals may be incomplete
- Advisory severity resolved for 120 of 137 advisories (lookup cap); the remainder are reported as unknown severity
- OpenSSF Scorecard did not return a usable result (exit code -9); skipping Scorecard checks

## What this report is

inspect.software measures public repositories against a single published, versioned methodology and reports the result as a 1–100 health index. Findings are signals derived from what is visible in public repository data — they are not a security audit, a warranty, or an endorsement, and no result depends on payment.

- Methodology, formulas and weights: https://inspect.software/methodology.md
- What a result does and does not claim: https://inspect.software/wiki/signals-not-warranties.md
- Rating bands: https://inspect.software/wiki/scoring-bands.md
- Corrections: https://inspect.software/contact.md
- Index of all documents: https://inspect.software/llms.txt
