Why I Tested Every Skill Instead of Just Listing Them
Most "best Claude skills" posts are just a GitHub README with a new headline slapped on top. Someone finds a repo, copies the description, ranks it by star count, and calls it a review. Nobody actually runs the thing.
That's a problem, because a skill that sounds useful in a bullet list can do nothing — or make things worse — on a real task. I wanted to know which skills actually change what Claude produces, not which ones have the best marketing copy.
So here's what I did instead: I installed each skill fresh, on a clean setup, and ran it against real coding tasks in my own repos. Not a demo repo built to make the skill look good. My actual code, with actual bugs and actual messy conventions.
For every skill worth your time, you'll get an install link, a plain description of what it's actually good and bad at, and a before/after — the same task, same prompt, skill off vs. skill on. You'll see the real output, not my summary of it. Some skills earned a spot on this list. A few didn't, and I'll tell you why.
What a Claude Skill Actually Is (And What It Isn't)
Before testing anything, it's worth being precise about what a "Claude Skill" is — because the term gets used loosely, and it's easy to confuse with three other things: subagents, MCP servers, and plugins.
A Skill is just a folder. Inside it, there's a SKILL.md file with instructions written in plain language, plus optional scripts or reference files. Claude doesn't run this folder all the time — it scans the short description in SKILL.md, decides whether the current task matches, and if so, loads the full instructions into context. If the task doesn't match, the skill sits unused. That's the whole mechanism. No magic, no persistent process running in the background.
Think of it like a laminated instruction card in a kitchen. It's not doing anything on its own — but when a specific situation comes up ("how do we plate this dish"), someone pulls the card out, reads it, and follows it. A Skill works the same way: dormant until Claude recognizes the moment it applies, then it shapes the response.
This is different from the other three pieces people lump together:
- MCP servers are a live connection to an external tool or API — a phone line Claude can pick up to check your database, hit GitHub's API, or read a file system in real time. It's not instructions, it's a working connection.
- Subagents are a full handoff. Claude delegates a task to a separate instance with its own context window, lets it work independently, and gets back a result. It's less "instruction card," more "coworker you hand a task to and wait on."
- Plugins are a bundle — one package that can include skills, custom commands, and MCP configs together, installed in one step. A plugin might contain three skills and an MCP server, all wired up at once.
Here's the short version, side by side:
| Type | What it is | Use it when... |
|---|---|---|
| Skill | Instructions Claude loads on demand | You want consistent behavior for a recurring task (code review style, test conventions) |
| Subagent | A delegated task with its own context | You want a task run independently without cluttering your main conversation |
| MCP server | A live connection to an external tool/API | Claude needs to actually read or act on something outside its training data |
| Plugin | A packaged bundle of the above | You want one-step setup for a workflow that needs multiple pieces working together |
Why this matters for this post: everything below is a Skill in the strict sense — a folder with instructions, not a live tool connection or a delegated agent. That's a meaningful limit. A Skill can tell Claude how to review a PR, but it can't fetch the PR itself. If a "skill" you find online needs an API key or a persistent connection, you're actually looking at an MCP server wearing a skill's name.
How I Actually Tested These (My Methodology)
Here's the part every listicle skips, and it's the part that actually matters.
I used the same repo and the same five coding tasks for every skill: fix a known bug, review a real PR, build a small UI component from a spec, write a test suite for an untested function, and debug a script that fails intermittently. Each task ran twice — once with the skill off, once with it on — same prompt, same model, same repo state.
"Worked" meant something specific, not a vibe. I looked for three things: fewer follow-up corrections needed to get a usable answer, correct output on the first try instead of the third, and alignment with real project conventions — actual docs, existing code style, real test patterns — instead of generic best practices that don't match the codebase.
If a skill made the output longer or more confident-sounding but didn't change whether the fix was correct or the test caught the right edge case, I counted that as no measurable effect. A few skills fell into exactly this bucket, and I say so plainly instead of padding the list to hit twelve.
If skill-on and skill-off produced the same fix with different formatting, that's not a win. That's noise.
One honest limitation: I tested on my own projects, which are mostly TypeScript and Python web apps. If you work in Rust, embedded systems, or a large legacy Java codebase, your results might differ — skills that lean on common web conventions may have less to work with elsewhere. I'll flag where I think a skill's usefulness is stack-specific rather than general, but I can't test every stack, and I won't pretend I did.
Installing a Skill in Under 2 Minutes (All 3 Ways)
Skills aren't complicated to install. The confusion online mostly comes from people mixing up where a skill lives, which changes who can use it. Here are all three install paths, and the actual folders involved.
Claude.ai (web or desktop app): Go to Settings > Capabilities > Skills. You can upload a skill folder directly or paste a URL to one. This installs it for your account only — it follows you across conversations but doesn't touch your local machine or your team.
Claude Code CLI: This is just a folder drop. Skills live in one of three places, and the location decides who benefits:
~/.claude/skills/— user-level. Available in every project you open on your machine. Good for personal habits, like how you like PRs described..claude/skills/inside a project — project-level. Checked into git, so anyone on the team gets it the moment they clone the repo. This is the one I use most, because coding conventions are usually project-specific, not personal.- Org-level, pushed through the admin console — every seat in your Claude Code org gets it automatically. Useful for company-wide standards (security review, PR format), overkill for anything experimental.
Via conversation: The laziest and honestly most reliable method — just tell Claude "install the skill at this URL" and it clones the repo and places it in the right folder for you. No terminal required.
Here's what a manual install actually looks like, so you know what you're dropping into place:
git clone https://github.com/example/pr-review-skill.git tmp-skill
mkdir -p ~/.claude/skills/pr-review
cp -r tmp-skill/* ~/.claude/skills/pr-review
rm -rf tmp-skill
# folder structure Claude expects:
~/.claude/skills/pr-review/
SKILL.md <- the instructions Claude reads
examples/ <- optional reference files
scripts/ <- optional helper scripts
That's it. No restart needed — Claude Code picks up new skills the next time you start a session. If nothing changes, check that SKILL.md actually exists at the top level; a skill nested one folder too deep is the most common install mistake I made while testing these.
The 12 Coding Skills Worth Installing (Tested, Ranked, With Before/After)
Ranked by how much they actually changed the output across my five tasks — not by how impressive they sound. A few near the bottom are real but narrow. I've said so.
1. Planning-with-files skill
Keeps a running plan.md that survives context resets, so Claude doesn't relearn what it already figured out.
Install: git clone .../persistent-planner ~/.claude/skills/planner
Tested on: the intermittent bug task, which took several back-and-forth turns.
Before: after the context compacted, Claude re-investigated a hypothesis it had already ruled out two turns earlier.
After: Claude read plan.md, saw "ruled out race condition in queue.js:40," and moved straight to the next hypothesis.
Limitation: it only helps if the skill actually checkpoints after each step. Some versions I found just create the file once and never update it — check that yours writes to it mid-task, not just at the start.
2. Debugging / root-cause skill
Forces reproduction and a stated hypothesis before Claude touches any code.
Install: git clone .../root-cause-debug ~/.claude/skills/debug
Tested on: the intermittent-failure script.
Before: Claude wrapped the flaky call in a retry loop — the test passed, the bug didn't.
After: it added targeted logging first, reproduced the failure, and found a shared cache being written from two threads.
Limitation: it's slower on purpose. For a genuinely trivial bug, the insistence on reproduction steps before fixing feels like overhead.
3. Code Reviewer skill
Reviews against your actual repo's patterns, not generic best practices.
Install: git clone .../pr-review-skill .claude/skills/pr-review
Tested on: a real open PR in my repo.
Before: "Looks good, consider adding tests."
After: "This bypasses the validateInput() helper used in every other handler in /api — see users.ts:44 for the pattern."
Limitation: only as good as what's actually in your repo. If conventions live in someone's head and not in code or docs, this skill has nothing to reference and defaults back to generic advice.
4. Test-writing skill
Matches existing test structure and fixtures instead of writing generic happy-path tests.
Install: git clone .../test-writer .claude/skills/tests
Tested on: an untested date-range overlap function.
Before: three happy-path tests, no edge cases.
After: added boundary-equality, single-point overlap, and reversed-range cases, using the same describe/it structure and fixture helper as the rest of the suite.
Limitation: it infers edge cases from the code, not from intent. If a function is subtly wrong, it may write a test that just confirms the wrong behavior.
5. API / docs lookup skill (Context7-style)
Pulls current library docs at request time instead of relying on training data.
Install: git clone .../live-docs-lookup ~/.claude/skills/docs
Tested on: the UI component build task, using a library that shipped a breaking change recently.
Before: Claude imported useFormState from react-dom — deprecated in React 19.
After: used useActionState with the current signature, matching the version actually in package.json.
Limitation: adds a few seconds of latency per lookup, and it's only as accurate as its match to your actual installed version — a stale lockfile can still fool it.
6. Codebase knowledge graph skill
Builds a map of how a repo's pieces actually connect before answering questions about it.
Install: git clone .../understand-repo .claude/skills/onboard
Tested on: asking Claude to explain the auth flow before the PR review task.
Before: guessed the middleware order from filenames, got it backwards.
After: traced the call chain across four files and correctly identified that token refresh happens in a request interceptor, not the login route.
Limitation: building the initial map takes real time — a few minutes on a ~40k-line repo. Not worth it on a small project.
7. Git / PR-description skill
Writes PR descriptions matching your repo's actual template.
Install: git clone .../pr-writer .claude/skills/pr-desc
Tested on: the bug-fix task.
Before: "Fixes bug in date handling."
After: full what/why/how-tested sections matching the repo's existing PR template, with the actual diff referenced.
Limitation: it's formatting, not substance. It didn't touch the code at all — useful, but the smallest lever on this list in terms of actual output quality.
8. Refactoring / large-codebase-navigation skill
Finds every usage of a piece of logic before changing it, instead of editing in place.
Install: git clone .../nav-refactor .claude/skills/refactor
Tested on: the UI component task, which needed logic extracted from a large shared file.
Before: edited inline, missed two other call sites using the same logic.
After: found all six usages via structured search, updated them consistently.
Limitation: it leans on grep-style search, not a real language server. On a genuinely huge monorepo it gets slow and can still miss dynamic references.
9. Database schema / safe migration skill
Writes migrations that account for live data, not just schema correctness.
Install: git clone .../safe-migrations .claude/skills/db
Tested on: I added a sixth task for this one, since none of my original five touch a database — adding a nullable column to a live table.
Before: a single migration adding a NOT NULL column directly — would lock the table and break existing rows.
After: a two-step migration (add nullable, backfill, then constrain), with the lock risk flagged explicitly.
Limitation: needs your actual schema or engine to reason correctly. If it assumes Postgres and you're on MySQL, the advice can be subtly wrong.
10. Security / vuln-scanning skill
Reviews specifically for security issues, overlapping with the code-review skill but narrower.
Install: git clone .../vuln-scan .claude/skills/security
Tested on: the PR review task.
Before: missed that a new endpoint didn't check ownership before querying by user-supplied ID.
After: flagged the possible IDOR (insecure direct object reference) and suggested an ownership check.
Limitation: also flagged three non-issues already handled at the middleware layer it couldn't see. You still need a human to triage the output.
11. Browser automation skill
5 Skills I Tried and Cut (And Why They Didn't Make the List)
Not everything I installed made the cut. Here's what I tried, what went wrong, and why I stopped using each one. If you're deciding what to skip, this should save you the setup time.
Diagram-generation skill. The idea is good: ask Claude to draw a Mermaid diagram of your architecture before you touch it. On a small repo, it worked fine. On my 40k-line test repo, it produced a diagram with 60+ nodes and no meaningful grouping — technically accurate, completely unreadable. I spent more time squinting at the output than I would have spent reading the code directly.
Pentesting skill. This one scans for a wider range of vulnerabilities than the narrower vuln-scan skill I kept (see #10 above). The problem is volume: on one file, it flagged 14 "issues," and 11 were false positives — things like flagging eval() usage in a test file that never runs in production. A security skill you don't trust is worse than no security skill, because you either ignore everything (missing the real ones) or triage everything manually (which defeats the point).
Over-engineered planning skill. This skill made Claude write a full task breakdown, dependency graph, and risk assessment before starting any coding task — including the bug-fix task, which was a two-line date-handling fix. For genuinely large tasks this kind of planning helps. For small ones, it added a multi-paragraph plan I had to read and approve before Claude touched a single line. The overhead cost more time than the plan ever saved me.
Humanizer-style skill. This is built for rewriting AI-sounding prose into more natural-sounding text — think blog posts, emails, marketing copy. It has nothing to do with code, and there's no version of a coding task where it's relevant. I only tried it because it showed up in a "top Claude skills" list with no context about what it's actually for. Worth naming here because you'll see it recommended in generic roundups that don't distinguish coding use cases from everything else.
Game-engine generator skill. This one is genuinely well-built — it scaffolds Unity or Godot project structures fast. But it only works if you're building a game. I tested it against my five non-game tasks and it either did nothing useful or actively tried to force a game-engine folder structure onto a web app. Not a bad skill. Just a skill for a completely different job than "coding" in general.
How to Stack Skills Without Confusing Claude
Once you have more than a handful of skills installed, a new problem shows up: which one does Claude actually use when two of them could apply?
Here's what's happening under the hood. Claude doesn't load every skill's full instructions into context all the time — that would burn tokens fast. Instead, it first reads just the name and description metadata for each installed skill. Based on that lightweight pass, it decides which skill (if any) is relevant to your request, and only then loads that skill's full content.
This means the description you write — or that a skill's author wrote — matters more than people expect. I ran into this directly: I had both the code-review skill (#1 on this list, not shown above) and the security/vuln-scan skill (#10) installed at the same time, with descriptions that both said something like "reviews code for issues." On a PR review task, Claude picked one at random-feeling and ignored the other. I only caught it because the output was missing the IDOR flag I'd seen in earlier tests.
The fix was rewriting the descriptions to be specific and non-overlapping: the code-review skill's description now says it checks for style, structure, and logic bugs; the security skill's says it checks specifically for auth, injection, and access-control issues. Once they stopped describing the same job, Claude used both when a task called for it.
Two more practical habits that helped:
- Keep a lean, project-scoped set. Instead of installing every skill globally, I install only what a given project needs at the project level (
.claude/skills/inside the repo, not your home directory). A Node API project doesn't need the game-engine skill sitting there as a candidate for Claude to consider. - Audit descriptions when something misfires. If Claude picks the wrong skill, the fix is almost never "delete a skill" — it's "make the descriptions more specific."
When You Should Just Write Your Own Skill Instead
Some of the skills on this list took real engineering to build. But the most useful skill I made during this whole test took five minutes, because it wasn't trying to solve a general problem — it was solving one specific, repeated annoyance in my own codebase.
Here's the situation: every time I asked Claude to add a new feature that needed to call an external service, it wrote a raw fetch() call. Reasonable default. Wrong for my codebase, though — we have an internal API client that handles auth headers, retries, and error formatting, and every raw fetch call was code I'd have to rewrite in review. That happened three times before I fixed it at the source instead of catching it every time.
The fix was a SKILL.md file with almost nothing in it:
---
name: use-internal-api-client
description: Use before writing any HTTP request to an external service. Enforces use of our internal API client instead of raw fetch/axios calls.
---
Before writing any code that makes an HTTP request:
1. Check src/lib/api-client.ts first. It wraps fetch with
our auth headers, retry logic, and error formatting.
2. Use `apiClient.get()` / `apiClient.post()` instead of
raw fetch() or axios calls.
3. If the client doesn't support what you need (e.g. a new
method or header), extend api-client.ts rather than
bypassing it.
Example of what NOT to do:
fetch('https://api.example.com/users')
Example of what to do instead:
apiClient.get('/users')
That's it. No install script, no external repo, no dependencies. I dropped it in .claude/skills/use-internal-api-client/SKILL.md in the project and it's been catching raw fetch calls ever since.
The case for writing your own over installing someone else's: the 500-line skills on this list are built to work across many codebases, which means they're built around assumptions — a testing framework, a schema convention, a file structure — that may or may not match yours. A 10-line skill about your specific internal convention doesn't need to generalize. It just needs to be true for your repo. If you find yourself correcting Claude for the same reason more than twice, that's the signal to write the skill instead of writing the correction again.
A Word on Trust: Don't Install Skills Blindly
Here's the part most "awesome skills" lists skip: a skill isn't just a text file telling Claude how to think. It can bundle scripts — shell commands, Python files, install steps — and those run with your permissions, on your machine. If you'd think twice before running a random script off GitHub, apply the same rule here. A skill is code you're agreeing to execute, not just a prompt you're agreeing to read.
This isn't a reason to panic. It's a reason to spend 60 seconds before you install anything from a source you don't already trust:
- Read the SKILL.md fully. Not skim it — read it. It's usually short. If you can't tell what it does after reading it, that's a red flag on its own.
- Check for embedded scripts. Look for shell commands,
curl | bashpatterns, or references to files outside the skill's own folder. If it wants to reach out to the internet or touch files it doesn't need to, ask why. - Check the repo's activity. Recent commits, more than one contributor, actual issues being closed — that's a maintained project. A single file uploaded once with no history is a coin flip.
- Prefer known sources over gists. A skill from a maintainer with a real repo and other people using it is a different risk profile than a one-off pastebin link someone posted in a Discord.
I skipped this check exactly once, on a skill that turned out to be harmless but poorly written. It wasn't malicious — just sloppy, with a script that shelled out to run tests in a way that silently failed. Nothing bad happened, but it cost me an afternoon of confused debugging before I found the source. The checklist above would have caught it in a minute.
Try This Today
Don't install all twelve. Pick one.
If you want the fastest payoff, start with the code-reviewer skill or the docs-lookup skill from earlier in this post — both give you a visible before/after in one run, so you'll know within minutes whether it's worth keeping.
Here's the whole plan:
- Copy the install command from the relevant section above.
- Drop it in
.claude/skills/inside a real project — not a test repo, the one you're actually working in. - Open a PR or bug you already have sitting open right now.
- Ask Claude to use the skill on it, and read the output like you'd read a real review — agree or disagree with each point.
That's it. No need to overhaul your setup or install a bundle. One skill, one real task, five minutes. You'll either see something worth catching that you missed, or you won't — and either way, you'll know more about whether skills are worth your time than you did before reading this.
I'm still testing new ones as they show up. If you want the next batch of before/after outputs, follow along at parkerjoseph.dev — I post the failures too, not just the ones that worked.
A Deeper Walkthrough: Running the Code-Reviewer Skill on a Real Pull Request
The before/after screenshots earlier in this post tell you what a skill produces. They don't tell you what it's like to actually run one on a Tuesday afternoon, on a PR that matters, when you're not sure if it's going to help or just add noise. So here's the whole process, start to finish, using the code-reviewer skill on a real pull request from one of my side projects. I'm including the parts that didn't go smoothly, because those are the parts that actually teach you something.
Step 1: Pick a task with a checkable outcome
Don't test a skill on a toy problem. You won't learn anything, because toy problems don't have the messy edge cases that separate a useful review from a generic one. I picked a PR that added rate limiting to an API route — real logic, a few edge cases around concurrent requests, and code I'd already reviewed myself so I had a baseline to compare against. That last part matters more than it sounds: if you don't already know what's wrong with the code, you have no way to judge whether the skill caught it or missed it.
If you're trying this yourself, use a PR you've already reviewed, or one with a bug you already know about. You're not testing whether the skill can write code — you're testing whether it can catch the same things you catch, in less time, or catch things you didn't.
Step 2: Install without breaking your existing setup
The install itself is one line — drop the skill folder into .claude/skills/ in the project root. Where it got slightly annoying was that I already had a CLAUDE.md file in that project with my own review preferences written into it (things like "flag any unhandled promise rejection" and "prefer early returns over nested conditionals"). I wasn't sure if the skill's instructions would override mine, get ignored, or stack on top.
The answer, after testing it: they stack, but not always cleanly. Claude treats the skill's SKILL.md as additional context, not a replacement for your project instructions. In practice this meant my "prefer early returns" preference showed up in the output alongside the skill's own checklist. But when the two gave conflicting guidance — my file said one thing about error handling, the skill's file implied another — Claude picked one without flagging the conflict. I only noticed because I read closely. If you have strong existing conventions written into a CLAUDE.md or .cursorrules file, read the skill's instructions side by side with yours before you run it, so you know where they might clash.
Step 3: Force the skill to trigger, because it might not on its own
Here's the part nobody mentions in the quick-start guides: skills don't always fire automatically. They're supposed to activate when Claude decides the task matches the skill's description, but that matching is fuzzier than you'd expect. On my first try, I just said "review this PR" and Claude gave me a normal review — no sign the skill had kicked in at all. No error, no explanation. It just didn't use it.
The fix was to be explicit: "Use the code-reviewer skill to review this PR." That worked every time. So did referencing the skill by name in my prompt when I described the task. If you install a skill and the output looks exactly like what you'd get without it, don't assume the skill is bad — check whether it actually ran first. You can usually tell because skill-driven output has a different shape: it follows the skill's own structure (a checklist, a severity ranking, specific section headers) instead of Claude's default review format. If your output doesn't have that shape, name the skill directly in your prompt and try again.
Step 4: Compare skill-on against skill-off, on the same code
This is the step that actually tells you whether a skill earns a permanent spot in your workflow, and it's the one most people skip. I ran the same PR through Claude twice — once with the skill invoked, once with a plain "review this code" prompt — and put the two outputs side by side.
The plain review was fine. It caught a missing null check and suggested renaming a variable. Useful, but generic — the kind of thing any decent reviewer catches on a first pass. The skill-driven review caught the same null check, plus something the plain review missed entirely: the rate limiter used an in-memory counter that would reset on every server restart, which meant it offered no protection in a multi-instance deployment. That's a real bug, and it's the kind of thing that's easy to miss because the code works fine in local testing and only breaks under conditions you don't see until production.
Why did the skill catch it and the plain prompt didn't? Because the skill's instructions explicitly told Claude to check for state that doesn't survive restarts or scale across instances — a specific, narrow instruction that a generic "review this code" prompt doesn't include. That's the actual value of a good skill: not that it makes Claude smarter, but that it gives Claude a checklist of specific things to look for that you'd otherwise have to type out yourself, every single time.
The trade-off showed up too. The skill-driven review took noticeably longer — it worked through its checklist methodically, which meant more back-and-forth before I got the final output. For a two-line fix, that overhead isn't worth it. For a PR with real logic and real risk, it was.
Step 5: Decide keep, tweak, or drop
After one real PR, I had enough information to make a call: keep it, but only invoke it explicitly on PRs above a certain size or risk level, not on every small change. That's the decision you're working toward with this whole exercise. You're not trying to find the "best" skill in the abstract — you're trying to find out where a specific skill earns its keep in your specific workflow, and where it's just overhead.
If a skill doesn't catch anything the plain prompt missed after two or three real tries, drop it. That's not a failure on your part — it just means that particular skill's checklist doesn't match the kind of bugs your codebase tends to produce. A skill built for reviewing Python data pipelines isn't going to add much if you're writing React components, even if it's well written. Match the skill to your actual failure modes, not to how good it sounds in its description.
FAQ
What's the actual difference between a Claude Skill and just writing a good system prompt?
A system prompt lives in one conversation and disappears when the context resets. A skill lives in a folder, gets versioned in your repo like any other file, and loads automatically whenever Claude decides the task matches. The practical difference is durability: you write the instructions once, commit them, and every teammate who clones the repo gets the same behavior without copying and pasting a prompt into their own setup. If you're the only person using a particular set of instructions, a saved prompt might be simpler. Once more than one person needs the same behavior, a skill is the better fit because it travels with the codebase instead of living in someone's notes.
Do skills work in Claude.ai, or only in Claude Code?
Skills are built for Claude Code and similar developer tools that read from a .claude/skills/ directory. The web version of Claude.ai doesn't currently look for that folder, so a skill sitting in your repo won't do anything if you paste code into the chat window instead of working through Claude Code. If you're doing most of your coding work in the browser, skills won't help you directly — you'd get more value from Projects or custom instructions, which serve a similar purpose in that context but work differently under the hood.
Do skills use up my context window?
Yes, but usually less than you'd think. The SKILL.md file gets loaded into context when the skill activates, so a bloated 2,000-word skill file costs you real tokens on every single run. This is one reason the well-written skills I tested tend to be short — most of the ones worth using are under 300 words. If you're building your own skill, or evaluating someone else's, treat length as a warning sign. A skill that needs three pages of instructions to explain a simple task is probably trying to do too much, and you'll pay for that in context on every invocation, not just the first one.
Can I run more than one skill at the same time?
Yes, and I do this regularly — the docs-lookup skill and the code-reviewer skill from earlier in this post work fine together, since one pulls reference material and the other applies review logic. Where it gets messy is when two skills give overlapping or conflicting instructions on the same kind of task, like two different code-style skills that disagree on formatting conventions. Claude doesn't reliably flag that conflict for you; it just picks one and moves on, and you might not notice until the output looks slightly off. If you're stacking skills, keep their responsibilities distinct — one for review, one for docs, one for testing — rather than installing multiple skills that all try to do the same job differently.
Does using a skill make responses slower?
Sometimes, yes. A skill that walks through a checklist — like the code-reviewer skill going through its list of specific checks — takes longer than a plain one-shot answer, because Claude is working through more explicit steps instead of just generating a response. For quick tasks, that overhead isn't worth it. For anything where accuracy matters more than speed — a security review, a database migration, a PR that touches payment logic — the extra time is a reasonable trade for catching something you'd otherwise miss. Match the skill's thoroughness to how much the task actually matters, not to whether a skill is available.
How do I update a skill after I've already installed it?
If you installed it by cloning or copying files into .claude/skills/, you update it the same way you'd update any other file in your repo: pull the latest version from the source, overwrite the old folder, commit the change. There's no separate update mechanism or version manager doing this for you — skills are just files, which is part of what makes them easy to inspect but also means you're responsible for tracking whether the source has changed. If you're using a skill from someone else's repo, it's worth checking back on it every month or so, the same way you'd check for updates to any dependency you didn't write yourself.
Do these skills work with models other than Claude?
No — the format is specific to Claude Code's skill system, and the way it triggers, loads context, and structures output doesn't carry over to other tools. Some of the underlying ideas do translate — a well-written checklist for code review is useful no matter what model reads it — but the file format, the auto-invocation behavior, and the folder structure are all Claude-specific. If you're working across multiple AI coding tools, you'll need a separate equivalent for each one; there's no single skill file that works everywhere yet.
Is it better to build my own skill instead of using one of the twelve from this post?
Depends on how specific your problem is. All twelve skills I tested were built for fairly common situations — code review, documentation lookup, that kind of thing. If your team has a genuinely unusual convention — a specific way you handle database migrations, or a house style for error messages that nobody else uses — writing your own skill for that is worth the twenty minutes it takes, because no public skill is going to know your internal rules. For anything more generic, start with an existing skill and tweak it rather than writing from scratch. It's faster, and you'll learn what a working SKILL.md looks like before you try to write your own from a blank page.
How do I cleanly remove a skill I don't want anymore?
Delete its folder from .claude/skills/ and commit the change. That's the whole process — there's no registry entry or config file elsewhere that needs cleaning up, since the skill's presence in that folder is the only thing that made it active in the first place. If you were referencing the skill by name in saved prompts or documentation, update those too, but the skill itself leaves no trace once the folder is gone.
Do skills replace things like linters, unit tests, or CI checks?
No, and I'd be careful about anyone who tells you otherwise. A linter catches the same category of mistake the same way, every time, with zero variability — that's exactly what you want for things like formatting or unused imports. A skill-driven review is more like a second pair of eyes: useful for catching things that require judgment, like the in-memory rate limiter bug from the walkthrough above, but not something you should rely on to be perfectly consistent run to run. Keep your linters and tests doing what they're good at. Use a skill for the layer above that — the judgment calls a static tool can't make.