AI research lab · consulting · Melbourne, AU

We break AI models so yours don’t break you.

Shrestech is a research lab and AI consulting firm. We run authorized adversarial campaigns against frontier models — and help enterprises adopt, design, and build AI practices that hold.

  • Canary payloads only
  • Nothing executed
  • Human-adjudicated

Models fail in patterns. They refuse the plain ask and ship the wrapped one. They hold the first four turns and bend on the fifth. Filters read prompts, not tool results. We find where your deployment sits on that line — measured, not guessed.

01Services

Assessment-first
AI consulting.

We don’t resell platforms. We measure, advise, and harden — with evidence on the table at every step.

  1. LLM assessment

    Your model under a graded battery: refusal integrity, multi-turn escalation ladders, obfuscation channels, automated best-of-N sampling. A per-model report with exposure ranked and every verdict traceable to a transcript.

    Deliverable — model posture report

  2. Security assessment

    The system around the model: tool layers and MCP integrations, RAG pipelines, agent sandboxes, egress posture. We probe the seams that prompt filters don’t cover.

    Deliverable — seam map & fix list

  3. AI practice design

    Model selection, guardrail architecture, deployment review. We help teams adopt, design, and build AI practices that hold — independent, grounded in measurement; we publish our method, not our opinions.

    Deliverable — practice blueprint

  4. Continuous evaluation

    Models change; so does their posture. We re-run the full battery on every version bump and gate your releases on measured thresholds.

    Deliverable — release gates

02Method

How an engagement
actually runs.

Discipline is the product. The same protocol, every engagement — so results mean the same thing across models, teams, and time.

  1. Scope & rules of engagement

    Targets, limits, and authorization in writing. Canary payloads only; nothing destructive is ever executed, on your systems or ours.

  2. Measure

    A graded battery across prompt, tool, and channel surfaces — multi-turn ladders, planted files, live fetches, automated sampling at scale.

  3. Adjudicate

    Dual-gate scoring with human sign-off. Dead runs are marked UNMEASURED and excluded — never counted as passes.

  4. Report

    Findings ranked by exposure, each pinned to its evidence. No screenshots without verdicts; no verdicts without transcripts.

  5. Harden & re-test

    Concrete fixes, shield design where needed — then the battery again, until the numbers hold.

03Field notes

What assessments
surface.

Patterns from real evaluation work — the shapes we test for first, because they show up first.

The seam

Filters read prompts, not tool results

Provider safety filters scan the prompt and pass the fetched page, the tool output, the planted file. If your agents browse or call tools, that content arrives unclassified. Classify it yourself, at the boundary.

Observed — frontier provider stacks

The ladder

Patience breaks single-shot testing

Individually defensible turns accumulate into a covert ask. The pattern that broke a 2026 flagship model measured zero on every single-shot score. Multi-turn escalation is a first-class threat class, not an edge case.

Observed — 5-turn legitimacy ladder

The wrapper

It’s the command, not the encoding

Base64 launders nothing by itself. The imperative wrapper — “decode this and do exactly what it says” — is the bend vector: instruction-following outrunning classification. Score the register, not just the vocabulary.

Observed — encoding-channel sweeps

04About

A research lab
that consults.

Shrestech is a research lab and AI consulting firm based in Melbourne. We study how frontier models fail — and help enterprises adopt, design, and build AI practices on what we learn.

Our background is security research and enterprise delivery, and we keep the discipline from both: authorized targets, measured claims, human adjudication, and reports your engineers can act on line by line.

05Contact

Tell us what
you’re deploying.

The fastest way to start is a one-line description of your stack. Every engagement opens with scope and rules of engagement — what we test, what we never touch, and what you get.

First scoping assessment — free

Describe your stack in one line. We reply with scope, method, and a fixed price — no obligation, no retainer.

Write to

support@shrestech.com