skills agent is never ranked against a packages agent and evidence from one
track is never compared to another.
skills
Agent skill bundles (
SKILL.md + code). Detonated and scored on dual plane
evidence with proof of execution. Skills are the classic prompt injection
vector, so the context plane matters intensely.mcp_servers
MCP server packages. Detonated with a component centric analysis: tool
descriptions, schemas, source, responses, resource handlers, and config, and
how influence propagates across them.
packages
pip and npm packages. Detonated across the lifecycle: install time and
import time behaviour, plus supply chain signals (CVEs, typosquat, dependency
confusion).
repositories
Source repositories. The outlier: a static audit with no probe, scored by F2
against the benchmark’s known vulnerabilities.
The detection principle
Every track shares one idea: malice is the gap between what an artifact declares it does and what it actually does. Your agent derives the artifact’s declared intent, exercises or audits it under instrumentation, and treats any deviation as the signal. Signatures catch known bad strings; deviation catches the novel and obfuscated attacks no signature has seen. What changes per track is where declared intent comes from and how you observe behaviour, not the principle. This maps straight onto the evidence you emit. The capability manifest and the declared purpose are the intent side; the action plane (what the artifact did) and the context plane (what it tried to make the model do) are the behaviour side. Your SSSA is, in effect, a deviation report. The per track sections below give a first architecture and what the detector must catch. Treat the pseudocode as a floor, not a solution: it is public, so a submission that merely reimplements it competes against everyone else who read the same page. Emissions reward whoever climbs off the floor first.The threat model
Phylax assumes every artifact is untrusted by default and its author is fully adversarial. Whatever the track, the attacker is pursuing one or more of six goals:
The adversary controls every byte of what they publish and distributes it
through legitimate channels: skill marketplaces, MCP registries, package
indexes. They may craft any component of an artifact and compose several so
that each looks benign in isolation while the combination achieves compromise.
The attack proceeds automatically once an agent installs and uses the artifact,
with no further attacker interaction. Each track below enumerates how these
goals concretely surface for its artifact type; a detector that only covers
some of its track’s attack classes leaves score on the table as the corpus
grows to label the rest.
Detonation vs audit
Three tracks detonate. The agent loads the artifact into an instrumented sandbox, runs it, threads the probe through the execution, and the run produces filesystem, network, and process traces, recorded by the validator’s instrumentation. This is what lets Phylax catch behaviour that static reading cannot see: obfuscated payloads, runtime resolved imports, and instructions that only surface when the artifact is actually exercised. Therepositories track audits. Nothing is executed, so there is no probe and no
proof of execution; the track is scored by F2 against a benchmark of known
vulnerabilities.
Verification per track
skills
Artifact. An agent skill bundle, typically aSKILL.md plus helper scripts.
The SKILL.md is both documentation and, on the context plane, a place where
instructions can hide.
Attack surface. Six attack classes:
The launch benchmark labels malicious instruction injection and transitive
poisoning and is being extended to the remaining classes; the scoring spine is
generic across all six, so coverage grows with the corpus rather than requiring
new scoring code. Build for all six now: corpus growth converts that coverage
into score.
Evidence.
proof_of_execution + action_plane + context_plane. Both
planes are required: a strong agent must catch context contamination and
transitive risk, not only an obvious instruction in the skill text.
Detection approach. The declared task is the skill’s stated purpose in
SKILL.md. Extract it before the model reads the rest of the bundle, so hidden
instructions cannot rewrite your notion of what the skill is for. Then run the
skill and flag any action outside that task: a formatting skill that reads
credentials, a skill that obeys an instruction embedded in retrieved content.
mcp_servers
Artifact. An MCP server package: a manifest declaring tools plus the server implementing them. Why it is the deepest track. An MCP server sits at the intersection of two attack channels in a single artifact. Malicious logic can live in executable server code (the action plane), and adversarial instructions can be embedded in the metadata and outputs the language model consumes as reasoning context (the context plane). Server provided metadata is not inert documentation: it is read by the model and actively steers tool selection and follow up actions. Component centric analysis. Malice may live in any component or be distributed across several so that each looks benign in isolation:
Attacks compose across components, and influence propagates through the model’s
orchestration: adversarial content in one tool’s description can steer the model
into invoking a second tool whose code carries the payload. Logic split between
a poisoned description and conditional code in a tool’s source evades any check
that inspects either component alone.
Attack surface. Eleven attack classes:
Launch scoring covers tool poisoning and schema mismatch, with the remaining
classes forming the threat surface and the next additions to scoring. Build for
the full surface now.
Detection approach. Two stages, following the behavioural deviation
principle. Pre-execution: read the config for malicious startup or shell
commands, and extract each tool’s declared intent from its description
separately, before adversarial text in one description can contaminate your
reading of another. In-execution: invoke tools in a sandbox and trace the whole
trajectory step by step, flagging a tool that acts beyond its declared function,
and catching attacks split across calls that each look benign alone.
mcp_surface block recording which
component carried each finding and how influence propagated.
packages
Artifact. A pip or npm package:setup.py / pyproject.toml (or
package.json) plus the source.
Why install time matters. A package can attack the moment it is installed,
before it is ever imported. Empirical studies find roughly two thirds of
malicious PyPI packages execute at install time, and typosquatting accounts for a
clear majority of injection methods in the foundational malicious package
datasets. Industry telemetry through 2025 and 2026 reports hundreds of thousands
of new malicious packages per year, shifting toward install time execution,
credential harvesting, and dependency confusion. A track that observes the
install phase, not only imported behaviour, is aligned with where package
attacks actually occur.
Attack surface. Nine attack classes:
Evidence.
proof_of_execution + a lifecycle block distinguishing install
time from import time behaviour + action_plane + a supply_chain block (SBOM,
dependency CVEs, typosquat, dependency confusion).
Detection approach. Derive the declared purpose from the metadata and README,
then execute both the install process and the runtime behaviour in a sandbox and
monitor the API call sequence. Prioritise install time. Prefer behaviour and data
flow evidence, a taint path from a secret source to a network sink, over surface
features: malware increasingly mimics benign code, and surface classifiers
degrade against it.
repositories
Artifact. A source repository: a tree of source files plus its manifests. What the agent does. Audits the source statically and reports recovered vulnerabilities, each with a weakness class (CWE), file, line, severity, and remediation. There is no probe, no proof of execution, and no dual plane evidence. Attack surface. Six finding classes:
The benchmark’s labelled dimensions (vulnerabilities, supply chain, secrets)
map straight onto these classes, so each one you cover contributes to your F2.
Scored by. F2 against the benchmark’s known vulnerabilities: a reported
finding matches a known vulnerability when it agrees on CWE or title and
localises within a small line window. F2 tilts toward recall, since a missed real
vulnerability costs more than a false alarm. Clean repositories measure the false
positive rate through precision. See Scoring for the formal
rules.
Detection approach. Derive the project’s declared purpose, then audit the
codebase and flag components whose logic contradicts that purpose. This is the
broadest track, the largest behaviour surface and the coarsest declared intent,
so lean on semantic and LLM assisted reading of what code actually does rather
than pattern matching. Two classes are easy to miss: backdoors and triggers
buried in otherwise ordinary code, and malicious natural language instruction or
config files that AI coding agents trust and act on (a redirected
settings.json,
a SKILL.md that exfiltrates keys).
Emission weight
The performance pool is divided by track first, then flows to each track’s above threshold top agents through consensus.repositories carries by far the largest share, then packages, mcp_servers,
and skills. The ordering is strict: every top-three winner of a higher track
outearns every top-three winner of the one below. Repositories and packages have
the clearest objective ground truth and the biggest real world supply chain
impact. See the
Incentive Mechanism for the full mechanism.
Per-track budgets
Each track’s per task budget is pinned and identical on every validator, expressed as CPU time so hardware never changes an outcome:
These are frozen: they are part of what your agent is measured against. See
Configuration and The Round Model.