• 2026Featured

    AMD-GAN

    Class-specific adaptive WGAN-GP framework that generates configurable synthetic network traffic to correct extreme class imbalance in intrusion detection.

    Each attack class gets its own dedicated generator, with adaptive configurations for minority classes (smaller batches, more epochs, stronger regularisation). Synthetic data quality is validated with Train-Synthetic-Test-Real protocols, and the resulting detectors are stress-tested against adversarial traffic. Validated on CIC-IDS2017.

    • GANs
    • NIDS
    • Class imbalance
    • TensorFlow
  • 2026Featured

    PACE Drift

    When to adapt: a cheap, unsupervised covariate-shift trigger for streaming intrusion detection (Proactive Adaptation via Covariate-shift Estimation).

    A reconstruction-error monitor over incoming traffic drives a dual Page-Hinkley trigger that fires before labelled feedback reveals a drop in accuracy. Across CIC-IDS2018 and UNSW-NB15, proactive triggering improves sustained F1 over reactive, error-based adaptation, at a cost below 0.5 ms per sample.

    • Concept drift
    • Streaming ML
    • NIDS
  • 2025Featured

    GenPot

    A generative honeypot that combines fine-tuned LLMs with OpenCanary and a FastAPI service to produce adaptive, realistic web and API interactions.

    Instead of serving static decoy responses, GenPot uses fine-tuned language models to generate plausible answers to attacker requests on the fly, keeping adversaries engaged while their behaviour is logged. It ships as a Docker image and is part of the CiberIA research initiative (INCIBE, Project C079/23).

    • Honeypots
    • LLMs
    • Cyber deception
    • FastAPI
    • Docker
  • 2026

    Binary Stylometry Attribution

    Authorship attribution from compiled binaries, for both benign code and malware, using Ghidra pseudocode and GraphCodeBERT.

    Source files are compiled at several optimisation levels (O0–O3), decompiled with Ghidra in headless mode and encoded with GraphCodeBERT to train a style classifier. Includes inference on new files and explainability reports (confusion matrices, confidence histograms, token importance).

    • Malware attribution
    • Reverse engineering
    • Transformers
  • 2026

    IoC Knowledge Engine

    Proof of concept of a knowledge-generation engine that correlates indicators of compromise and telemetry with MITRE ATT&CK using semantic search and a local LLM.

    Network traces, malware classifications and honeypot logs are matched against ATT&CK techniques via FAISS and sentence-transformers; a local LLM then produces a structured analysis of TTPs, risks and recommendations. Ships with simulated scenarios to try end to end.

    • Threat intelligence
    • MITRE ATT&CK
    • LLMs
    • FAISS
  • 2025

    HoneyV Malware Dataset

    Honeypot-based malware analysis laboratory with a curated collection of 65+ real-world malware families across 8 categories.

    Automated workflow from sample acquisition and honeypot deployment to hybrid (static and dynamic) forensic analysis in isolated virtual machines. Each family includes the sample plus hashes and VirusTotal references. Contains live malware — handle only in isolated environments.

    • Malware analysis
    • Honeypots
    • Datasets
  • 2025

    Maintainable AI-driven Threat Detection

    Comparative study of 112 intrusion-detection experiments across four datasets, preprocessing techniques and models, plus a modular detection framework.

    Evaluates KNN, Random Forest, Linear SVC, LightGBM, XGBoost, stacking and neural networks on CIC-IDS2017/2018/2019 and UNSW-NB15, with PCA vs. Top-K feature selection and with/without SMOTE. First activity of the CiberIA project.

    • NIDS
    • Benchmarking
    • Machine learning
  • 2025

    Multimodal Malware Family Classification

    End-to-end pipeline that classifies malware families from static, dynamic and visual (binary-to-image) features and compares early, late and expert-selection fusion at scale.

    Covers data preparation, feature extraction, CNN training on byte images, classical models (SVM, Random Forest, LightGBM, XGBoost, voting ensembles), explainability scripts and statistical comparison between fusion strategies.

    • Malware
    • Multimodal fusion
    • CNNs
    • XAI