Program · ActivePost-training

Project
North Star

A long-horizon program to build a model that finds unknown flaws, chains them the way a real operator would, and proves impact under lab control — so defenders see the attack path before an adversary walks it.

Parameters

0B

Stage

03 / 06

Weights

Closed

Release

TBD

Mission

Give defenders the attacker's foresight. Find the unknown flaw, prove the path, close it first.

01 · Find

Search for the flaws nobody has catalogued yet — the ones a signature database cannot contain.

02 · Prove

Demonstrate real impact with bounded proof-of-exploit, so severity is measured rather than argued.

03 · Close

Ship the detection, the mitigation, and the patch guidance the same day the finding lands.

Thesis

Defensive security is bottlenecked by the shortage of people who can think like an attacker. That shortage is a modelling problem.

Most security tooling enumerates. It lists findings, ranks them by a static score, and leaves the reasoning to a human. The scarce skill is the next step — connecting unremarkable findings into a path that actually compromises something that matters.

North Star exists to build and bound that capability in the open research sense: measured, restricted, and pointed at defense.

Polaris

Post-training

The first model of the North Star program.

Parameters
300B
Specialization
Offensive security
Core skills
0-day · chaining · PoE
Modality
Text + tool use
Stage
Post-training
Weights
Closed
Capabilities under study

Zero-day discovery, exploit chaining, and proof-of-exploit

Capability

Zero-day discovery

Polaris reads source, binaries, and protocol state machines looking for the class of flaw a scanner cannot name: logic inversions, lifetime confusion, trust boundaries that only exist by convention. Candidate findings are triaged against reachability before a human ever sees them.

Illustrative trace · lab environment only

01

Zero-day discovery

Hypothesis-driven search for previously unknown flaws in code, binaries, and protocols — ranked by reachability, not by pattern count.

02

Exploit chaining

Composing low-severity primitives into full attack paths, with the privilege gained at each hop made explicit so defenders can cut the cheapest link.

03

Exploit development & PoE

Minimal, bounded proof-of-exploit artifacts built inside instrumented ranges to prove impact — paired with detections and patch guidance.

04

Defensive translation

Every offensive result is converted into detection logic, mitigations, and prioritized remediation. Offense is the method; defense is the product.

Lifecycle tracker

Where Polaris is today

Hover a star to inspect the stage
01Pretraining02Mid-training03Post-training04Safety evaluation05Red-team hardening06Limited release
Stage 03 · Post-training

Preference optimization, tool use, and refusal calibration. Polaris is here today.

In progress
Safety

Commitments

An offensive-security model is a dual-use artifact. These constraints are part of the program design, not a policy added at release.

  • No open weights for offensive-specialized models.
  • Capability thresholds are defined before training, not after results arrive.
  • All agentic evaluation runs inside isolated, instrumented lab environments.
  • Proof-of-exploit artifacts are minimal, non-weaponized, and never released publicly.
  • Access is limited to vetted defensive partners under written scope.
  • Findings are disclosed to affected parties before publication.