Marketing Foundation
Compact two-skill starter: clarify positioning and choose a lead magnet. Use Marketing Launch for the broader eight-skill go-to-market workflow.
The tunnel uses the same cards as the catalogue. Browse only as deep as needed — or load a broad bundle immediately.
SEO, Sales, Agents or another broad area → one bundle call → work.
Read-only access to published skills. Default 8, maximum 10 skills / 120,000 characters.
Compact two-skill starter: clarify positioning and choose a lead magnet. Use Marketing Launch for the broader eight-skill go-to-market workflow.
Build an evidence-led marketing plan from ICP and competition through positioning, campaigns, growth and measurement.
Diagnose architecture and context, plan agent-team responsibilities, then organize project context and session handoffs. Memory and cost-runtime reviews remain outside this pack.
Review the journey from landing page and lead capture through registration, first value and transparent upgrades.
Plan a campaign, draft its channel content and review the work against actual brand guidance.
Prioritize an editorial roadmap and plan how to launch and distribute it across suitable channels.
Understand customer needs, compare competitors and plan a community around real member value.
Choose a relevant lead magnet, then draft a permission-based nurture journey with entry, suppression and exit rules.
Define the API contract, then plan how to observe its latency, failures and retries. Guidance and checklist; no production changes.
Profile a dataset, choose and interpret statistical methods, then validate calculations and conclusions before sharing.
Original workflow for marketing prompt evaluation, version history and governance.
---
name: "prompt-engineer-toolkit"
description: "Turns marketing prompts into tested, versioned production assets: A/B prompt evaluation against structured test cases, immutable prompt version history with diffs, ready-to-use marketing prompt templates (ad copy, email campaigns, social posts, landing pages, SEO meta), and an LLM-governance playbook for marketing teams (claim discipline, disclosure rules, human-review gates). Use when a marketing team relies on AI-generated content and needs prompt quality to be measurable and safe — or when the user mentions 'prompt engineering,' 'improve my prompts,' 'prompt templates,' 'prompt versioning,' 'AI content workflow,' or 'AI governance for marketing.'"
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: marketing
updated: 2026-03-06
---
# Prompt Engineer Toolkit
## Overview
Use this skill to move prompts from ad-hoc drafts to production assets with repeatable testing, versioning, and regression safety. It emphasizes measurable quality over intuition. Apply it when launching a new LLM feature that needs reliable outputs, when prompt quality degrades after model or instruction changes, when multiple team members edit prompts and need history/diffs, when you need evidence-based prompt choice for production rollout, or when you want consistent prompt governance across environments.
## Core Capabilities
- A/B prompt evaluation against structured test cases
- Quantitative scoring for adherence, relevance, and safety checks
- Prompt version tracking with immutable history and changelog
- Prompt diffs to review behavior-impacting edits
- Reusable prompt templates and selection guidance
- Regression-friendly workflows for model/prompt updates
## Key Workflows
### 1. Run Prompt A/B Test
Prepare JSON test cases and run:
```bash
python3 scripts/prompt_tester.py \
--prompt-a-file prompts/a.txt \
--prompt-b-file prompts/b.txt \
--cases-file testcases.json \
--runner-cmd 'my-llm-cli --prompt {prompt} --input {input}' \
--format text
```
Input can also come from stdin/`--input` JSON payload.
### 2. Choose Winner With Evidence
The tester scores outputs per case and aggregates:
- expected content coverage
- forbidden content violations
- regex/format compliance
- output length sanity
Use the higher-scoring prompt as candidate baseline, then run regression suite.
### 3. Version Prompts
```bash
# Add version
python3 scripts/prompt_versioner.py add \
--name support_classifier \
--prompt-file prompts/support_v3.txt \
--author alice
# Diff versions
python3 scripts/prompt_versioner.py diff --name support_classifier --from-version 2 --to-version 3
# Changelog
python3 scripts/prompt_versioner.py changelog --name support_classifier
```
### 4. Regression Loop
1. Store baseline version.
2. Propose prompt edits.
3. Re-run A/B test.
4. Promote only if score and safety constraints improve.
## Script Interfaces
- `python3 scripts/prompt_tester.py --help`
- Reads prompts/cases from stdin or `--input`
- Optional external runner command
- Emits text or JSON metrics
- `python3 scripts/prompt_versioner.py --help`
- Manages prompt history (`add`, `list`, `diff`, `changelog`)
- Stores metadata and content snapshots locally
## Pitfalls, Best Practices & Review Checklist
**Avoid these mistakes:**
1. Picking prompts from single-case outputs — use a realistic, edge-case-rich test suite.
2. Changing prompt and model simultaneously — always isolate variables.
3. Missing `must_not_contain` (forbidden-content) checks in evaluation criteria.
4. Editing prompts without version metadata, author, or change rationale.
5. Skipping semantic diffs before deploying a new prompt version.
6. Optimizing one benchmark while harming edge cases — track the full suite.
7. Model swap without rerunning the baseline A/B suite.
**Before promoting any prompt, confirm:**
- [ ] Task intent is explicit and unambiguous.
- [ ] Output schema/format is explicit.
- [ ] Safety and exclusion constraints are explicit.
- [ ] No contradictory instructions.
- [ ] No unnecessary verbosity tokens.
- [ ] A/B score improves and violation count stays at zero.
## References
- [references/prompt-templates.md](references/prompt-templates.md) — 6 production marketing templates (ad copy, email sequence, social repurposing, landing sections, SEO meta, brand-voice rewrite) plus generic building blocks; each written to be graded by `prompt_tester.py`
- [references/technique-guide.md](references/technique-guide.md) — technique-selection table for marketing tasks + the LLM-governance stack for marketing teams (claim discipline, disclosure rules, data boundaries, human-review gates)
- [references/evaluation-rubric.md](references/evaluation-rubric.md) — mechanical scoring weights, acceptance gates, marketing quality dimensions, test-suite design, and eval anti-patterns
- [README.md](README.md)
## Evaluation Design
Each test case should define:
- `input`: realistic production-like input
- `expected_contains`: required markers/content
- `forbidden_contains`: disallowed phrases or unsafe content
- `expected_regex`: required structural patterns
This enables deterministic grading across prompt variants.
## Versioning Policy
- Use semantic prompt identifiers per feature (`support_classifier`, `ad_copy_shortform`).
- Record author + change note for every revision.
- Never overwrite historical versions.
- Diff before promoting a new prompt to production.
## Rollout Strategy
1. Create baseline prompt version.
2. Propose candidate prompt.
3. Run A/B suite against same cases.
4. Promote only if winner improves average and keeps violation count at zero.
5. Track post-release feedback and feed new failure cases back into test suite.
Original workflow for marketing prompt evaluation, version history and governance.
The complete original workflow, with source attribution and the limitations below.
Confirm the task, check actual dependencies, then apply the relevant original instructions within authorized scope.
Prompt use case, test cases and the separate upstream runner/versioner/reference package.
Published after an obvious-danger screen under the user-requested policy, not a functional test. prompt_tester.py, prompt_versioner.py, reference guides and README are not supplied or executed. Any external runner and local writes require task-scoped permission. Regex/substring scores are not semantic or safety proof; no runner is certified.
Use Marketing Prompt Toolkit for [TASK]. Ask for missing inputs: Prompt use case, test cases and the separate upstream runner/versioner/reference package. Check the declared dependencies and limitations before execution. Return the original output structure with known facts, assumptions and unresolved requirements separated.
Claiming that unavailable source dependencies are bundled or that generated outputs, integrations or model behavior have been tested. Automatic account access, file changes, publishing or paid operations outside the user task.
M11 added German routing, task inputs and explicit source-package limitations; the original author remains separate from M11 curation. Published after an obvious-danger screen under the user-requested policy, not a functional test. prompt_tester.py, prompt_versioner.py, reference guides and README are not supplied or executed. Any external runner and local writes require task-scoped permission. Regex/substring scores are not semantic or safety proof; no runner is certified.
MIT License Copyright (c) 2025 Alireza Rezvani Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
Copy the text below, then paste it into your chat.