AISSA / mail systems research

Open source · Rspamd + Postfix · MIT

What can an LLM
add to a mail filter?

AISSA explores selective content analysis inside an existing mail stack, with bounded waiting and measurable results.

Keep the familiar checks. Select messages worth a closer look. Measure whether the model adds useful evidence.

Experimental. Start in observation mode.

AI Sample Spam AnalyzerLocal Ollama backendOptional live scoringIndependent Redis collection

01 / Integration

Fit analysis into the mail stack.

Existing Rspamd checks run first. AISSA adds a separate selection and analysis step, while Rspamd keeps control of the final action.

  1. 01

    Select

    Lua applies score, symbol, URL and reputation criteria, followed by sampling and size limits.

  2. 02

    Analyze

    An authenticated loopback bridge extracts bounded MIME text and calls a local Ollama model.

  3. 03

    Contribute

    Observation returns after submission. Optional live scoring waits within the remaining scan budget.

  4. 04

    Collect

    The controller records results independently. AI suspicion and independently confirmed abuse stay separate.

Timeouts remain part of the design.

A late or failed result adds status information, not a ham verdict. The live wait reserves scan completion time. This is not a hard deadline for the entire SMTP/Milter chain.

02 / Measured behavior

Publish the misses, too.

Manual observations with CPU-only Qwen 2.5 1.5B. These cases show behavior, not a production accuracy rate or throughput benchmark.

Selected model calls recorded on 3 October 2026
MessageVerdictModel timeObservation
Explicit password / 2FA demandphishing15.128 sCold request; live points contributed to rejection
Same demand, warm repeatphishing2.523 sMinimal load and prompt processing time
Changed meeting arrangementham4.195 s456 of 536 input tokens reused from cache
Suspicious donation offerham5.400 sMissed suspicious content; confidence 1.0

Model time excludes other filters and SMTP overhead. Test-server phishing points were +5; the repository example uses +3. Read the setup and field notes.

Prompt caching helps latency.

A warm model reused the shared input prefix even when the mail text changed. Retention improves that opportunity; cache hits are not guaranteed.

Confidence does not prove correctness.

The model gave its missed donation offer a confidence of 1.0. That number is a self-assessment and does not multiply the configured score.

03 / Current boundaries

Know what is being evaluated.

Implemented today

  • Local Ollama inference and validated JSON
  • Observation and optional bounded live scoring
  • Queue, rate and retained-result limits
  • Deduplicated Redis counters and result collection

Still to establish

  • Reliable decisions across real abuse and legitimate mail
  • Selection beyond a simple total-score range
  • Model comparisons and optional external API backends
  • Expiration and cancellation of stale model jobs

Attachments, malware, OCR and fetched web pages are outside the analyzer. Queues and results are held in RAM. Model work may continue after a live timeout. Hostile mail can cause invented evidence or prompt-injection failures.

04 / Participate

Bring mail expertise.
Bring counterexamples.

Useful contributions are independently labeled, privacy-reviewed cases; model comparisons; and integration reports from other Rspamd versions and SMTP paths.

Report false positives and missed abuse alongside latency, hardware and prompt version. Do not post private mail or credentials.