For Agents · MCPNo local installation · No M11 login
M11Capability Engineby Patrick Moser-Brillowski
Curated AI capabilities

Find a skill.
Get to work.

Useful AI skills from strong open sources, cleaned up for discovery, task fit and direct use.

All skills

◎265k ★GitHub

Multi-Agent Orchestration

Coordinates multi-agent work with clear owners, work items, evidence, and merge gates.

Agents · Orchestration→
·265k ★GitHub

AI Context Window Audit

Audits Claude Code context overhead and recommends ways to reduce unnecessary loaded content.

Research · Context→
·265k ★GitHub

AI Agent Architecture Audit

Diagnoses agent-system failures across prompts, memory, tools, wrappers, and output delivery.

Research · Agent Architecture→
≡265k ★GitHub

AI Skill Discovery

Searches local and external skill sources for existing matches before a new skill is created.

Knowledge Work · Skill Discovery→
↗51k ★GitHub

Lead Magnet Strategy

Plans lead magnets around audience needs, buyer stage, capture approach, distribution, and measurement.

Marketing · Lead Generation→
↗27k ★GitHub

Ideal Customer Profile

Defines an evidence-based ideal customer profile from research, customer behavior and jobs to be done.

Marketing · Customer Research→
↗27k ★GitHub

Go-to-Market Strategy

Builds a launch plan connecting target segments, channels, messaging, milestones and measurable outcomes.

Marketing · Go-to-Market→
↗27k ★GitHub

Marketing Campaign Ideas

Generates five campaign concepts with audience messages, channel choices and testable engagement hypotheses.

Marketing · Campaigns→
↗27k ★GitHub

Product Growth Loops

Evaluates product-led growth loops and outlines measurable experiments for sharing, collaboration and referrals.

Marketing · Growth→
↗27k ★GitHub

Competitor Analysis

Compares competitors using cited evidence and identifies differentiation opportunities and research gaps.

Marketing · Market Research→
↗27k ★GitHub

North Star Metric

Defines one customer-value metric and supporting input metrics with clear measurement assumptions.

Marketing · Measurement→
↗27k ★GitHub

Product Positioning

Develops differentiated product positioning ideas with audience fit, rationale, and supporting messages.

Marketing · Positioning→
·0 ★GitHub

Product Vision

Draft and compare product vision statements grounded in company values and customer needs.

Business · Product Strategy→
·0 ★GitHub

Go-to-Market Motion Selection

Compare seven acquisition approaches and prioritize a practical go-to-market plan.

Business · Go-to-Market→
·0 ★GitHub

Customer Feedback and JTBD Analysis

Synthesize supplied feedback into evidence-backed themes, jobs to be done and improvement priorities.

Business · Customer Research→
·0 ★GitHub

PESTLE Market Environment Analysis

Map external political, economic, social, technological, legal and environmental factors for a business decision.

Business · Product Strategy→
·0 ★GitHub

Customer Journey Mapping

Map customer touchpoints and friction from awareness through advocacy.

Business · Customer Research→
·0 ★GitHub

Ansoff Growth Options

Compare growth options across existing and new products and markets.

Business · Product Strategy→
·0 ★GitHub

Data Analysis Validation

Review methodology, calculations and conclusions before sharing an analysis.

Data & Analytics · Data Analysis→
·0 ★GitHub

Dataset Profiling

Profile a dataset and identify quality issues and useful follow-up analyses.

Data & Analytics · Data Analysis→
·0 ★GitHub

Statistical Analysis Guidance

Choose descriptive statistics and hypothesis tests while making assumptions and uncertainty explicit.

Data & Analytics · Data Analysis→
↗0 ★GitHub

Programmatic SEO Planning

Plan useful SEO pages at scale with a data strategy, templates and twelve complete playbooks.

Marketing · SEO→
↗0 ★GitHub

Landing Page and Form Conversion Review

Review marketing pages and forms, prioritize friction fixes and design measurable experiments.

Marketing · Conversion Optimization→
↗0 ★GitHub

Paywall and Upgrade Planning

Plan transparent in-product upgrade prompts and experiments after users experience value.

Marketing · Conversion Optimization→
↗0 ★GitHub

Signup and Registration Review

Review account creation and trial signup friction while preserving necessary security and consent controls.

Marketing · Conversion Optimization→
↗0 ★GitHub

User Onboarding and Activation

Plan the first useful product experience, activation milestones and measurable onboarding experiments.

Marketing · Conversion Optimization→
↗0 ★GitHub

Popup and Modal Planning

Design dismissible, accessible conversion overlays with honest offers and measurable frequency rules.

Marketing · Conversion Optimization→
↗0 ★GitHub

Lifecycle Email Sequence Planning

Draft a complete lifecycle email sequence with timing, branching, exits and suppression rules.

Marketing · Email Marketing→
↗0 ★GitHub

Marketing Campaign Planning

Build a campaign brief with audience, messages, channel choices, calendar, dependencies and measurement.

Marketing · Campaign Planning→
↗0 ★GitHub

Marketing Content Drafting

Draft channel-specific marketing content using clear structures, evidence and calls to action.

Marketing · Content Marketing→
↗0 ★GitHub

Brand Voice and Content Review

Review drafts against supplied brand guidance and propose specific, prioritized revisions.

Marketing · Brand Strategy→
↗0 ★GitHub

Marketing Performance Reporting

Turn supplied campaign or channel metrics into a traceable report with comparisons and testable recommendations.

Marketing · Marketing Analytics→
◇0 ★GitHub

Sales Company Research

Research a company or partner and produce a sourced fit hypothesis and draft outreach approach.

Sales · Company Research→
◇0 ★GitHub

Company and Contact Enrichment

Resolve company and contact records with field-level evidence, visible coverage limits and explicit match criteria.

Sales · Sales Intelligence→
↗0 ★GitHub

Website Information Architecture

Plan page hierarchy, navigation, stable URL patterns and useful internal links for a website.

Marketing · Website Architecture→
For Agents · Remote MCP

Let your chat find the right skill.

No local installation. No M11 login. Connect once. Broad task? Load 5–10 relevant skills and go. Precise task? Narrow through category, topic and tags.

Read only5–10 bundleCategory → Topic → Tags
For Agents · MCP

Your task.
The right skill.

The tunnel uses the same cards as the catalogue. Browse only as deep as needed — or load a broad bundle immediately.

No local installationNo M11 loginRead-only
Broad task
Load 5–10 and go

SEO, Sales, Agents or another broad area → one bundle call → work.

MCP endpoint: https://skills.m11.ch/mcp

Read-only access to published skills. Default 8, maximum 10 skills / 120,000 characters.

Task Packs

Marketing Foundation

Clarify positioning, then turn it into a lead-generation asset.

Marketing Launch

Build an evidence-led marketing plan from ICP and competition through positioning, campaigns, growth and measurement.

Agent Audit Essentials

Diagnose architecture and context, then plan agent-team responsibilities. Memory and cost-runtime reviews are outside this pack.

Conversion and Activation Review

Review the journey from landing page and lead capture through registration, first value and transparent upgrades.

Campaign Content and Brand Review

Plan a campaign, draft its channel content and review the work against actual brand guidance.

Choose an area

01 · Category
What the MCP returns at this step
02 · Research · Agent Architecture

AI Agent Architecture Audit

Diagnoses agent-system failures across prompts, memory, tools, wrappers, and output delivery.

debuggingagent-architectureauditpromptsmemorytoolswrappersreliability
Agent Architecture Audit · Original SKILL.md
---
name: agent-architecture-audit
description: Full-stack diagnostic for agent and LLM applications. Audits the 12-layer agent stack for wrapper regression, memory pollution, tool discipline failures, hidden repair loops, and rendering corruption. Produces severity-ranked findings with code-first fixes. Essential for developers building agent applications, autonomous loops, or any LLM-powered feature. Use when an agent or LLM feature misbehaves and the failing layer is unknown, or before shipping an agent stack.
metadata:
  origin: oh-my-agent-check
tools: Read, Write, Edit, Bash, Grep, Glob
---

# Agent Architecture Audit

A diagnostic workflow for agent systems that hide failures behind wrapper layers, stale memory, retry loops, or transport/rendering mutations.

## When to Activate

**MANDATORY for:**
- Releasing any agent or LLM-powered application to production
- Shipping features with tool calling, memory, or multi-step workflows
- Agent behavior degrades after adding wrapper layers
- User reports "the agent is getting worse" or "tools are flaky"
- Same model works in playground but breaks inside your wrapper
- Debugging agent behavior for more than 15 minutes without finding root cause

**Especially critical when:**
- You've added new prompt layers, tool definitions, or memory systems
- Different agents in your system behave inconsistently
- The model was fine yesterday but is hallucinating today
- You suspect hidden repair/retry loops silently mutating responses

**Do not use for:**
- General code debugging — use `agent-introspection-debugging`
- Code review — use language-specific reviewer agents
- Security scanning — use `security-review` or `security-review/scan`
- Agent performance benchmarking — use `agent-eval`
- Writing new features — use the appropriate workflow skill

## The 12-Layer Stack

Every agent system has these layers. Any of them can corrupt the answer:

| # | Layer | What Goes Wrong |
|---|-------|----------------|
| 1 | System prompt | Conflicting instructions, instruction bloat |
| 2 | Session history | Stale context injection from previous turns |
| 3 | Long-term memory | Pollution across sessions, old topics in new conversations |
| 4 | Distillation | Compressed artifacts re-entering as pseudo-facts |
| 5 | Active recall | Redundant re-summary layers wasting context |
| 6 | Tool selection | Wrong tool routing, model skips required tools |
| 7 | Tool execution | Hallucinated execution — claims to call but doesn't |
| 8 | Tool interpretation | Misread or ignored tool output |
| 9 | Answer shaping | Format corruption in final response |
| 10 | Platform rendering | Transport-layer mutation (UI, API, CLI mutates valid answers) |
| 11 | Hidden repair loops | Silent fallback/retry agents running second LLM pass |
| 12 | Persistence | Expired state or cached artifacts reused as live evidence |

## Common Failure Patterns

### 1. Wrapper Regression

The base model produces correct answers, but the wrapper layers make it worse.

**Symptoms:**
- Model works fine in playground or direct API call, breaks in your agent
- Added a new prompt layer, existing behavior degraded
- Agent sounds confident but is confidently wrong
- "It was working before the last update"

### 2. Memory Contamination

Old topics leak into new conversations through history, memory retrieval, or distillation.

**Symptoms:**
- Agent brings up unrelated past topics
- User corrections don't stick (old memory overwrites new)
- Same-session artifacts re-enter as pseudo-facts
- Memory grows without bound, degrading response quality over time

### 3. Tool Discipline Failure

Tools are declared in the prompt but not enforced in code. The model skips them or hallucinates execution.

**Symptoms:**
- "Must use tool X" in prompt, but model answers without calling it
- Tool results look correct but were never actually executed
- Different tools fight over the same responsibility
- Model uses tool when it shouldn't, or skips it when it must

### 4. Rendering/Transport Corruption

The agent's internal answer is correct, but the platform layer mutates it during delivery.

**Symptoms:**
- Logs show correct answer, user sees broken output
- Markdown rendering, JSON parsing, or streaming fragments corrupt valid responses
- Hidden fallback agent quietly replaces the answer before delivery
- Output differs between terminal and UI

### 5. Hidden Agent Layers

Silent repair, retry, summarization, or recall agents run without explicit contracts.

**Symptoms:**
- Output changes between internal generation and user delivery
- "Auto-fix" loops run a second LLM pass the user doesn't know about
- Multiple agents modify the same output without coordination
- Answers get "smoothed" or "corrected" by invisible layers

## Audit Workflow

### Phase 1: Scope

Define what you're auditing:

- **Target system** — what agent application?
- **Entrypoints** — how do users interact with it?
- **Model stack** — which LLM(s) and providers?
- **Symptoms** — what does the user report?
- **Time window** — when did it start?
- **Layers to audit** — which of the 12 layers apply?

### Phase 2: Evidence Collection

Gather evidence from the codebase:

- **Source code** — agent loop, tool router, memory admission, prompt assembly
- **Logs** — historical session traces, tool call records
- **Config** — prompt templates, tool schemas, provider settings
- **Memory files** — SOPs, knowledge bases, session archives

Use `rg` to search for anti-patterns:

```bash
# Tool requirements expressed only in prompt text (not code)
rg "must.*tool|必须.*工具|required.*call" --type md

# Tool execution without validation
rg "tool_call|toolCall|tool_use" --type py --type ts

# Hidden LLM calls outside main agent loop
rg "completion|chat\.create|messages\.create|llm\.invoke"

# Memory admission without user-correction priority
rg "memory.*admit|long.*term.*update|persist.*memory" --type py --type ts

# Fallback loops that run additional LLM calls
rg "fallback|retry.*llm|repair.*prompt|re-?prompt" --type py --type ts

# Silent output mutation
rg "mutate|rewrite.*response|transform.*output|shap" --type py --type ts
```

### Phase 3: Failure Mapping

For each finding, document:

- **Symptom** — what the user sees
- **Mechanism** — how the wrapper causes it
- **Source layer** — which of the 12 layers
- **Root cause** — the deepest cause
- **Evidence** — file:line or log:row reference
- **Confidence** — 0.0 to 1.0

### Phase 4: Fix Strategy

Default fix order (code-first, not prompt-first):

1. **Code-gate tool requirements** — enforce in code, not just prompt text
2. **Remove or narrow hidden repair agents** — make fallback explicit with contracts
3. **Reduce context duplication** — same info through prompt + history + memory + distillation
4. **Tighten memory admission** — user corrections > agent assertions
5. **Tighten distillation triggers** — don't compress what shouldn't be compressed
6. **Reduce rendering mutation** — pass-through, don't transform
7. **Convert to typed JSON envelopes** — structured internal flow, not freeform prose

## Severity Model

| Level | Meaning | Action |
|-------|---------|--------|
| `critical` | Agent can confidently produce wrong operational behavior | Fix before next release |
| `high` | Agent frequently degrades correctness or stability | Fix this sprint |
| `medium` | Correctness usually survives but output is fragile or wasteful | Plan for next cycle |
| `low` | Mostly cosmetic or maintainability issues | Backlog |

## Output Format

Present findings to the user in this order:

1. **Severity-ranked findings** (most critical first)
2. **Architecture diagnosis** (which layer corrupted what, and why)
3. **Ordered fix plan** (code-first, not prompt-first)

Do not lead with compliments or summaries. If the system is broken, say so directly.

## Quick Diagnostic Questions

When auditing an agent system, answer these:

| # | Question | If Yes → |
|---|----------|----------|
| 1 | Can the model skip a required tool and still answer? | Tool not code-gated |
| 2 | Does old conversation content appear in new turns? | Memory contamination |
| 3 | Is the same info in system prompt AND memory AND history? | Context duplication |
| 4 | Does the platform run a second LLM pass before delivery? | Hidden repair loop |
| 5 | Does the output differ between internal generation and user delivery? | Rendering corruption |
| 6 | Are "must use tool X" rules only in prompt text? | Tool discipline failure |
| 7 | Can the agent's own monologue become persistent memory? | Memory poisoning |

## Anti-Patterns to Avoid

- Avoid blaming the model before falsifying wrapper-layer regressions.
- Avoid blaming memory without showing the contamination path.
- Do not let a clean current state erase a dirty historical incident.
- Do not treat markdown prose as a trustworthy internal protocol.
- Do not accept "must use tool" in prompt text when code never enforces it.
- Keep findings direct, evidence-backed, and severity-ranked.

## Report Schema

Audits should produce structured reports following this shape:

```json
{
  "schema_version": "ecc.agent-architecture-audit.report.v1",
  "executive_verdict": {
    "overall_health": "high_risk",
    "primary_failure_mode": "string",
    "most_urgent_fix": "string"
  },
  "scope": {
    "target_name": "string",
    "model_stack": ["string"],
    "layers_to_audit": ["string"]
  },
  "findings": [
    {
      "severity": "critical|high|medium|low",
      "title": "string",
      "mechanism": "string",
      "source_layer": "string",
      "root_cause": "string",
      "evidence_refs": ["file:line"],
      "confidence": 0.0,
      "recommended_fix": "string"
    }
  ],
  "ordered_fix_plan": [
    { "order": 1, "goal": "string", "why_now": "string", "expected_effect": "string" }
  ]
}
```

## Related Skills

- `agent-introspection-debugging` — Debug agent runtime failures (loops, timeouts, state errors)
- `agent-eval` — Benchmark agent performance head-to-head
- `security-review` — Security audit for code and configuration
- `autonomous-agent-harness` — Set up autonomous agent operations
- `agent-harness-construction` — Build agent harnesses from scratch

When to use

Use when an agent or LLM feature behaves inconsistently, tool use is unreliable, or a system needs a pre-release architecture review. It is not a general code-debugging or security-scanning workflow.

What you get

• Severity-ranked findings
• Architecture diagnosis
• Ordered fix plan

How it works

1. Define the system, symptoms, and audit scope.
2. Collect code, configuration, logs, and memory evidence.
3. Map failure causes and propose code-first fixes.

Requirements

Requires access to the relevant application code, logs, configuration, and memory artifacts; shell and search tools such as `rg` are used in the examples. No separate reference files are specified. Optional upstream related skills are not included in this catalog: agent-introspection-debugging, agent-eval, security-review, autonomous-agent-harness, agent-harness-construction. M11 relations are curated follow-up guidance, not equivalent replacements. Do not claim these external skills have been loaded.

Starting prompt

Use the skill below to investigate [AGENT FAILURE]. Ask for relevant code, logs and configuration. Separate evidence from hypotheses and return an ordered fix plan before making changes.

Not for

A full security certification or a substitute for evidence from code and runtime logs.

What M11 added

M11 added diagnostic routing, evidence requirements, clear exclusions and stable packaging around the unchanged agent-architecture audit.

Original authorship remains with affaan-m/ECC · Original source ↗
Original license & copyright
MIT License

Copyright (c) 2026 Affaan Mustafa

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.

Related skills & prerequisites

ai-context-window-audit · together — Inspect context overhead when diagnosing architecture failures.

multi-agent-orchestration · next — Translate the diagnosis into bounded agent ownership and handoffs.

Source SHA-256: 64f57e232c3533877403703bc95b3c75df9de23f657899a306f13c703009ff57
Snapshot checked: 2026-09-22T15:05:00.000Z
For Agents · MCP

Your next task. One connection.

No local installation · No M11 login