Marshorn Technologies · Research MVP

Prove agent security against ground truth—not the model’s word.

Gjallar runs attacker and defender agents in isolated, versioned labs and checks the environment itself. Linux and Windows AD fixtures are operational internally. No public SaaS yet.

Company

AI agents get tools. Safety claims stay anecdotal.

Marshorn Technologies Co., Ltd. builds Gjallar for security product, agent-governance, and research teams who need replayable proof—not a one-off pentest or a leaderboard screenshot.

Problem

Unverifiable agent security

Most claims rest on the model saying it succeeded, a demo, or a test that cannot be rerun when the model or policy changes.

Customers

Security, platform, research

Teams that ship agent controls or detection, and labs that need oracle-confirmed, fixture-versioned evaluation.

Industry

Evaluation, not pentesting

We ask a narrower question: do agent controls still hold under repeatable attacker/defender pressure?

Product

Isolated lab. Independent oracle. Replayable fixture.

Dual-agent episodes

Attacker and defender act in the same environment. Neither sees the other’s reasoning.

Ground-truth oracles

Success is a domain object, a ticket, a file—never “the agent thinks it worked.”

Frozen baselines

Versioned Linux and three-node Windows AD labs. Same digest, same starting state.

Research MVP. Companion Gjallar Plugin (OpenAI Build Week provenance, not endorsement) is a separate governance codebase.

Status

What is frozen, and what is next.

Operational

Linux baseline

Frozen enterprise-slice fixture with oracle-checked chain nodes.

Operational

Windows AD lab

DC, member, workstation. Identity chain and trusted oracle frozen. Live runs are isolated and self-hosted.

Next

Comparison + defender sight

Public multi-model report, then audit-based defender visibility on the AD chain.

Difference

Not a pentest scanner. Not a model leaderboard.

Autonomous pentesting

Find paths

Broad discovery is the product. We do not claim broader coverage.

Benchmarks

Rank models

Useful scores. They rarely connect a failure back to an operational control.

Gjallar

Test governance

Replay the same lab. Oracle-verify the outcome. Improve the control.

Team

Marshorn Technologies

Scott Fang

Scott Fang

Founder & Lead Researcher

Lab fixtures, oracles, and the dual-agent loop. Focus: post-compromise Windows AD evaluation.

Lin-Chun-Yu (Miranda)

Lin-Chun-Yu (Miranda)

AI Security Researcher

Responsible for AI security and the Linux track at Marshorn / Gjallar.

Oscar Yuh

Oscar Yuh

Audit Consultant

Responsible for audit and governance oversight at Marshorn.

Contact

Scoped research pilots.

One security or research partner, plus infrastructure for repeatable multi-VM evaluation.

Email
scott.fang@marshorn.com

Entity
Marshorn Technologies Co., Ltd.

Stage
Pre-seed research MVP. No hosted SaaS.