NLP · Cybersecurity · Temporal evaluation
Do NLP vulnerability classifiers survive time?
A study of CVE-to-CWE classification under chronological distribution shift, with controlled preprocessing, explicit exclusions, and auditable evidence.
Publication-ready research results: Pending. Historical benchmark results are superseded because their protocol allowed label-selection and description leakage.
Research questions
- How does a conventional IID estimate compare with performance on later publication years?
- Which text representations and preprocessing choices generalize best?
- How do language change, entity shortcuts, and ambiguous labels relate to observed errors?
Dataset
English NVD API 2.0 descriptions with exactly one specific CWE. The top classes are defined using 2018–2023 alone. Later out-of-vocabulary records and duplicate exclusions are reported explicitly.
Verified final dataset: Pending collection and validation.
NLP pipeline
NVD records → explicit filtering → frozen labels → leakage validation → training → validation selection → separate held-out evaluations → analysis
Feature vocabularies are fitted on training text. Test and future results do not select hyperparameters.
Tokenization & preprocessing
Raw text, lowercase, punctuation, stopwords, stemming, lemmatization, security-aware tokens, and CVE/version/IP masking are controlled ablations. Vendor and product masking uses a lexicon derived from training CPE records.
Aggressive preprocessing is a hypothesis to test, not an assumed improvement.
Classical models
Multinomial Naive Bayes, Logistic Regression, and Linear SVM with Bag of Words, word TF-IDF, word bigrams, and character n-grams. Macro-F1 is the selection metric.
Sentence embeddings
MiniLM (all-MiniLM-L6-v2) encodes descriptions; a logistic classifier is tuned on validation. Embeddings are cached by content and model revision.
Execution: Pending
DistilBERT
Sequence classification with deterministic label mappings, validation Macro-F1 checkpoint selection, and early stopping. Model weights are published only after successful training.
Execution: Pending
Temporal drift
Vocabulary Jaccard overlap, new-token rate, Jensen–Shannon divergence, and TF-IDF and sentence-embedding centroid distances compare historical language with each evaluation year. Associations with performance do not establish causation.
Final drift findings: Pending
Experiment registry
Each run records dataset hash, code commit, cutoff, configuration, seed, hardware, runtime, metrics, and artifact locations. Earlier runs remain available for audit.
- Implemented
- Code exists.
- Executed
- The stage actually ran.
- Verified
- Its generated artifacts were inspected and validated.
- Pending
- Unexecuted or blocked.
Results
Metrics below are loaded from generated research artifacts. No missing result is estimated or filled in.
Final model comparison: Pending
Statistical analysis
Bootstrap confidence intervals use the frozen class vocabulary. Paired comparisons require identical CVE IDs and labels. McNemar tests compare correctness, while paired bootstrap comparisons estimate Macro-F1 differences.
Final statistical findings: Pending
Explainability
Linear coefficients describe class-associated features. Transformer token occlusion measures changes after masking a token; it can create unnatural inputs and does not identify causal mechanisms. Raw attention is not treated as an explanation.
Verified explanations: Pending
Error analysis
Future errors are exported for human review with gold/predicted labels and candidate tags for terminology changes, confusable classes, description length, and annotation ambiguity. Automated tagging does not count as manual review.
Human review: Pending
Reproducibility
python scripts/reproduce.py --config configs/final.yamlCollection resumes from complete windows. Critical leakage failures stop experiments. Expensive stages require appropriate resources; stage status is reported explicitly.
Paper
Manuscript completion is Pending verified experiments. The study asks whether CVE-to-CWE classifiers that appear strong under conventional evaluation remain reliable under chronological shift.
Manuscript sources ↗Resources
Hugging Face dataset, model, and Space publication: Pending. Links will appear only after repository and public-access verification.
About
Arun Kumar Gharami · DriftCVE-NLP research project.
Research software under the MIT license. NVD source records retain their source provenance and limitations. This classifier is not an authoritative vulnerability assessment.