DriftCVE-NLPVulnerability language across time

NLP · Cybersecurity · Temporal evaluation

Do NLP vulnerability classifiers survive time?

A study of CVE-to-CWE classification under chronological distribution shift, with controlled preprocessing, explicit exclusions, and auditable evidence.

Publication-ready research results: Pending. Historical benchmark results are superseded because their protocol allowed label-selection and description leakage.

Research questions

  1. How does a conventional IID estimate compare with performance on later publication years?
  2. Which text representations and preprocessing choices generalize best?
  3. How do language change, entity shortcuts, and ambiguous labels relate to observed errors?

Dataset

English NVD API 2.0 descriptions with exactly one specific CWE. The top classes are defined using 2018–2023 alone. Later out-of-vocabulary records and duplicate exclusions are reported explicitly.

Verified final dataset: Pending collection and validation.

2018–2023Train
2024Validation
2025Held-out test
Partial 2026Out-of-time · exact cutoff in manifest

NLP pipeline

NVD records → explicit filtering → frozen labels → leakage validation → training → validation selection → separate held-out evaluations → analysis

Feature vocabularies are fitted on training text. Test and future results do not select hyperparameters.

Tokenization & preprocessing

Raw text, lowercase, punctuation, stopwords, stemming, lemmatization, security-aware tokens, and CVE/version/IP masking are controlled ablations. Vendor and product masking uses a lexicon derived from training CPE records.

Aggressive preprocessing is a hypothesis to test, not an assumed improvement.

Classical models

Multinomial Naive Bayes, Logistic Regression, and Linear SVM with Bag of Words, word TF-IDF, word bigrams, and character n-grams. Macro-F1 is the selection metric.

Sentence embeddings

MiniLM (all-MiniLM-L6-v2) encodes descriptions; a logistic classifier is tuned on validation. Embeddings are cached by content and model revision.

Execution: Pending

DistilBERT

Sequence classification with deterministic label mappings, validation Macro-F1 checkpoint selection, and early stopping. Model weights are published only after successful training.

Execution: Pending

Temporal drift

Vocabulary Jaccard overlap, new-token rate, Jensen–Shannon divergence, and TF-IDF and sentence-embedding centroid distances compare historical language with each evaluation year. Associations with performance do not establish causation.

Final drift findings: Pending

Experiment registry

Each run records dataset hash, code commit, cutoff, configuration, seed, hardware, runtime, metrics, and artifact locations. Earlier runs remain available for audit.

Implemented
Code exists.
Executed
The stage actually ran.
Verified
Its generated artifacts were inspected and validated.
Pending
Unexecuted or blocked.

Results

Metrics below are loaded from generated research artifacts. No missing result is estimated or filled in.

Final model comparison: Pending

Statistical analysis

Bootstrap confidence intervals use the frozen class vocabulary. Paired comparisons require identical CVE IDs and labels. McNemar tests compare correctness, while paired bootstrap comparisons estimate Macro-F1 differences.

Final statistical findings: Pending

Explainability

Linear coefficients describe class-associated features. Transformer token occlusion measures changes after masking a token; it can create unnatural inputs and does not identify causal mechanisms. Raw attention is not treated as an explanation.

Verified explanations: Pending

Error analysis

Future errors are exported for human review with gold/predicted labels and candidate tags for terminology changes, confusable classes, description length, and annotation ambiguity. Automated tagging does not count as manual review.

Human review: Pending

Reproducibility

python scripts/reproduce.py --config configs/final.yaml

Collection resumes from complete windows. Critical leakage failures stop experiments. Expensive stages require appropriate resources; stage status is reported explicitly.

Paper

Manuscript completion is Pending verified experiments. The study asks whether CVE-to-CWE classifiers that appear strong under conventional evaluation remain reliable under chronological shift.

Manuscript sources ↗

Resources

Hugging Face dataset, model, and Space publication: Pending. Links will appear only after repository and public-access verification.

About

Arun Kumar Gharami · DriftCVE-NLP research project.

Research software under the MIT license. NVD source records retain their source provenance and limitations. This classifier is not an authoritative vulnerability assessment.