Praca z agentami kodującymi

Od ticketa do zmergowanego PR‑a w sześciu skillach

Skille i plik AGENTS.md, na których opieram codzienną pracę z agentami. Krótkie, konkretne i gotowe do skopiowania do własnego repozytorium.

Flow

Każdy skill obsługuje jeden etap pracy nad zmianą. Agent wykonuje robotę, a decyzje — co naprawiamy, jak i kiedy mergujemy — zostają po mojej stronie.

  1. Stałe zasady AGENTS.md

    Fundament. Agent czyta ten plik na starcie każdej sesji, więc nie powtarzam w promptach, jak uruchamiać testy, czego nie wolno dotykać i jak ma wyglądać commit.

  2. Analiza zgłoszenia /investigate-ticket

    Agent nie wierzy ticketowi na słowo: rozbija go na twierdzenia, każde sprawdza w kodzie, odtwarza błąd testem i oddaje raport z opcjami naprawy. Przed implementacją się zatrzymuje.

  3. Doprecyzowanie planu /grill-me

    Zanim powstanie kod, agent przepytuje mnie z planu — jedno pytanie naraz, każde z rekomendowaną odpowiedzią — aż nie zostanie żadna nierozstrzygnięta decyzja.

  4. Implementacja i sprawdzenie w aplikacji /local-app-login

    Agent pisze zmianę, loguje się do lokalnej aplikacji na odpowiednie konto i sprawdza efekt w przeglądarce. Dla zmian frontendowych AGENTS.md wymaga screenshotów w trzech rozdzielczościach.

  5. Niezależny review /autoreview

    Zmianę ogląda drugi model, który nie brał udziału w jej pisaniu. Jego uwagi to porady do zweryfikowania, a nie polecenia do wykonania.

  6. Pilnowanie PR-a /babysit

    Agent obserwuje CI i komentarze z review, poprawia prawdziwe problemy, odpowiada na fałszywe alarmy i pilnuje rebase'a — aż PR będzie gotowy do merge'a. Sam nie merguje.

  7. Merge i zamknięcie ticketa /babysit-merge

    Wariant na koniec: po doprowadzeniu PR-a do gotowości agent robi squash-merge i przesuwa ticket na Done. Tylko na moje wyraźne wywołanie i tylko dla kodu, który widziałem.

Skille

Skill to katalog z plikiem SKILL.md: nagłówek mówi agentowi, kiedy po niego sięgnąć, a treść — co robić. Poniżej opis po polsku i oryginalna treść każdego pliku.

investigate-ticket

analiza

Bada ticket albo zgłoszony problem i proponuje naprawę do omówienia, zanim powstanie jakakolwiek implementacja.

  • Rozbija ticket na osobne twierdzenia (objawy, warunki, podejrzane przyczyny, oczekiwania) i weryfikuje każde w kodzie, testach i historii.
  • Idzie ścieżką wywołań od punktu wejścia do linii, w której zachowanie się psuje — nie zatrzymuje się na pierwszej prawdopodobnej przyczynie.
  • Dla błędów pisze i uruchamia test regresyjny, który failuje na obecnym kodzie.
  • Jeśli potrzebuje danych z bazy, prosi o nie, zamiast odpytywać ją samodzielnie.
  • Wynikiem jest raport w pięciu sekcjach: co jest nie tak, twierdzenia z ticketa, reprodukcja, przyczyna źródłowa, opcje naprawy z rekomendacją.
investigate-ticket/SKILL.md
raw
---
name: investigate-ticket
description: Investigate a ticket or reported problem and propose a fix for discussion before implementation.
---

Read the ticket and any supplied context, then thoroughly investigate the relevant code.
Split the ticket into separate claims (symptoms, conditions, suspected causes, expectations) and verify each one against the code, tests, and history instead of accepting the ticket's framing.
Follow the call path from the entry point to the line where the behaviour goes wrong, including every module it delegates to (workers, queues, frontend state, external APIs). Do not stop at the first plausible cause.
Before proposing a change to shared code, list its existing callers and check how each would be affected.
For bugs, write and run regression tests that reproduce the problem and fail against the current implementation. If a claim needs production or local database data to confirm, ask for it instead of querying.
Stop before implementing the fix until we have discussed and agreed on the approach.

## Report

The report is the deliverable. A developer who has not seen the code must be able to judge from it alone whether the proposed implementation makes sense, without redoing the investigation. A short summary is not an acceptable report. Write it in the user's language and use these sections:

1. **What is wrong** — plain-language explanation: what the user sees, who is affected, under which conditions, and why it happens. Explain the mechanism in words first, then point at the code.
2. **Ticket claims** — every claim from the ticket marked as confirmed, not confirmed, contradicted, or not verifiable, each with its evidence (`file:line`, test output, commit). Add relevant findings the ticket does not mention.
3. **Reproduction** — the test file and name, the command to run it, the relevant failing output, and the conditions it depends on (data, plan, flags, timing). If the problem did not reproduce, list what was tried and what that rules out. For non-bug tickets, show the current behaviour instead.
4. **Root cause** — the call path from entry point to faulty line with `file:line` at each hop, why that line is wrong, and when it was introduced if history explains it.
5. **Fix options** — for each option:
   - the concrete changes per file and function, including migrations, API or contract changes, and frontend changes;
   - why it removes the root cause rather than the symptom;
   - the affected callers and behaviour changes elsewhere;
   - what happens to existing data already in the bad state and whether it needs a backfill or cleanup;
   - risks and edge cases;
   - the tests to add or change;
   - the relative size of the change.

   End with a recommendation and its reasoning, and list the open questions that need a decision before implementation.

grill-me

planowanie

Agent bezlitośnie przepytuje mnie z planu lub projektu, dopóki nie dojdziemy do wspólnego rozumienia każdej gałęzi drzewa decyzji.

  • Pytania zadaje pojedynczo.
  • Do każdego pytania dodaje swoją rekomendowaną odpowiedź.
  • Jeśli odpowiedź da się znaleźć w kodzie, szuka jej tam zamiast pytać.
grill-me/SKILL.md
raw
---
name: grill-me
description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when the user wants to stress-test a plan, get grilled on their design, needs clarification before implementation, or mentions "grill me" or the "grill-me" skill.
---

# Grill Me

Interview the user relentlessly about every aspect of the plan until you reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one by one.

For each question, provide your recommended answer.

Ask the questions one at a time.

If a question can be answered by exploring the codebase, explore the codebase instead.

local-app-login

testowanieprzykład skilla projektowego

Pozwala agentowi wybrać konto i organizację w lokalnej aplikacji i zalogować się na nie w przeglądarce — wyłącznie w środowisku deweloperskim.

  • Jedno polecenie listuje dostępne konta, drugie generuje link logowania dla wybranej pary klient–organizacja.
  • Komendy mix app.accounts i mix app.login to własne taski projektu, więc tego skilla nie da się użyć 1:1. Do skopiowania jest wzorzec: daj agentowi prosty, opisany sposób wejścia do lokalnej aplikacji.
  • Skill odsyła do --help jako źródła prawdy, zamiast dublować dokumentację poleceń.
local-app-login/SKILL.md
raw
---
name: local-app-login
description: Select a local app account and organization and log in using development-only Mix commands. Use for local app browser testing, not the admin panel or production.
---

In the current checkout's app directory, list available accounts:

```bash
mix app.accounts
```

Choose a customer–organization pair suitable for the task and generate a login link:

```bash
mix app.login --customer <integer> --organization <guid>
```

Open the printed URL in the browser and verify the selected account and organization.
The app must already be running.

For command options and output details, use `mix app.accounts --help` or
`mix app.login --help`. Help is authoritative.

Login persists in localStorage across reloads and switches the account in all tabs
on the same origin (scheme, host and port). Worktrees on different ports are isolated.

autoreview

reviewopenclaw · MIT

Ustrukturyzowany code review wykonywany przez niezależny model. Jedyny skill w kolekcji, który nie jest mój — pochodzi z repozytorium openclaw/agent-skills i jest tu w niezmienionej postaci.

  • Domyślnie recenzentem jest model OpenAI uruchamiany przez Codex; Claude wchodzi do gry, gdy go wskażę albo gdy Codex jest niedostępny.
  • Zakres review wybiera się jawnie: lokalne zmiany, gałąź względem bazy albo pojedynczy commit.
  • Domyślny próg to tylko P0, czyli realne blokery; szerszy review włącza się flagą --max-priority.
  • Znaleziska to porady do zweryfikowania, a nie instrukcje do ślepego wdrożenia.
  • Skill to nie tylko SKILL.md: zawiera skrypt pomocniczy, dokumentację wyników i testy, dlatego zamiast wklejać treść odsyłam do źródła.
skąd wziąć
git clone https://github.com/openclaw/agent-skills.git
cp -R agent-skills/skills/autoreview .agents/skills/autoreview

babysit

pull request

Pilnuje PR-a, aż będzie gotowy do merge'a: obserwuje CI, komentarze i nierozwiązane wątki review jednocześnie.

  • Push nigdy nie kończy zadania — po każdym agent wraca do obserwowania.
  • Poprawki trzymają się pierwotnego zakresu PR-a. Sugestie nowych funkcji wymagają mojej zgody, nawet jeśli są sensowne.
  • Każde znalezisko jest sprawdzane na aktualnym kodzie. Fałszywy alarm dostaje jednozdaniową odpowiedź i zamknięcie wątku; uwagi subiektywne i architektoniczne trafiają do mnie.
  • Przed każdym pushem: fetch, rebase na origin/main i komplet checków na kodzie po rebase'ie.
  • Potwierdzone flaki infrastruktury ponawia najwyżej dwa razy, potem zgłasza blokadę.
  • Gotowość ogłasza dopiero po świeżym sprawdzeniu CI, wymaganych review, botów recenzujących i mergeowalności. Nie merguje, jeśli o to nie poproszę.
babysit/SKILL.md
raw
---
name: babysit
description: Monitor a PR until ready to merge. Use when asked to babysit a PR, watch CI, or keep handling review feedback.
---

Poll CI, comments, and unresolved review threads together; handle review feedback without waiting for CI to finish. A push is never the end of the task; resume watching.
Keep review fixes within the PR's original scope. Treat suggestions to add features or unrelated improvements critically; even if worthwhile, get explicit user approval before implementing them.
Verify findings against the latest code. Fix real issues within scope, including cheap nits; run relevant tests, commit, push.
Before every push, fetch origin and rebase onto origin/main; resolve conflicts and run all required checks on the rebased code before pushing.
For false positives, reply with a one-line reason and resolve the thread. Escalate valid findings you would rather defer, and subjective or architectural feedback, instead of dismissing them.
Resolve threads only after verifying they are addressed; outdated does not mean resolved.
Retry confirmed infra flakes up to twice, then report the blocker; if a check also fails on main, report it as a blocker.
Prefix comments you post with "<model slug> on behalf of <name>".
Before declaring readiness, freshly confirm that CI passes, required reviews are satisfied, no automated reviewer is still running (check their status comments and reactions), the PR is mergeable, and no comments or threads remain unaddressed. Check recent merged PRs for the review bots this repo uses; a bot that has not posted yet has not started, which is not a pass.
Stop when ready, merged, closed, or blocked on a human decision; briefly report the outcome or blocker.
Do not merge unless explicitly asked.

babysit-merge

pull requesttylko ręcznie

To samo co babysit, a na końcu squash-merge i przesunięcie ticketa w Jirze na Done.

  • Agent nigdy nie uruchamia tego skilla z własnej inicjatywy — w nagłówku jest disable-model-invocation: true.
  • Moje wywołanie zatwierdza kod wyłącznie w stanie z tamtej chwili. Jeśli kod PR-a zmieni się z jakiegokolwiek powodu — także przez poprawki samego agenta — merge'a nie będzie, dopóki nie wywołam skilla ponownie.
  • Jakakolwiek wątpliwość co do kodu też wstrzymuje merge.
babysit-merge/SKILL.md
raw
---
name: babysit-merge
description: Babysit a PR, then squash-merge it and move its Jira ticket to Done. Only when the user explicitly invokes /babysit-merge; never invoke on your own.
disable-model-invocation: true
---

Babysit the PR, and once it is ready, squash-merge it and move its Jira ticket to Done.
The user's invocation approves only the code as it was at that moment, so do not merge if the PR's code has changed since, for any reason including your own fixes, or if you have any open question or doubt about the code; report it and wait for the user to invoke this skill again.

AGENTS.md

Skille opisują pojedyncze zadania, a AGENTS.md — zasady obowiązujące zawsze.

Granice bezpieczeństwa

Żadnego dotykania środowisk zdalnych bez wyraźnej zgody. Na infrastrukturze agent tylko czyta; polecenia zmieniające stan może mi najwyżej zaproponować.

Kontrakt zamiast defensywy

Żadnych nowych null-checków, fallbacków i warunków „na wszelki wypadek”. Jeśli nie wiadomo, czy wartość może być pusta, agent sprawdza typy, schemat i miejsca wywołań.

Bez kodu legacy

Stare ścieżki się zastępuje, a nie utrzymuje obok nowych. Niezmienniki wymusza się u źródła.

Checki przed commitem

Lista checków pochodzi z konfiguracji CI, ale lokalnie uruchamia się warianty deweloperskie. Wszystko musi przejść — niezależnie od tego, kto zepsuł.

Commity i komentarze tłumaczą „dlaczego”

Każdy commit ma treść wyjaśniającą intencję zmiany. Komentarze w kodzie nie powtarzają tego, co widać w kodzie.

Review to sugestie

Każdą uwagę z review agent weryfikuje w kodzie i odrzuca te błędne, zamiast je posłusznie wdrażać.

Bezpieczne migracje

Bez kolumn z wartością domyślną i bez UPDATE-ów na całej tabeli w istniejących tabelach — mogą zablokować tabelę na czas migracji.

Testy nie zaglądają w konfigurację

Testy nie sprawdzają wartości ani kształtu konfiguracji produkcyjnej, a zmiany samej konfiguracji nie dodają testów.

Frontend sprawdzany wzrokiem

Zmiany frontendowe agent ogląda na screenshotach w rozdzielczości mobilnej, tabletowej i desktopowej i dołącza je do opisu PR-a.

AGENTS.md
raw
# AI Agent Guidelines

This document provides guidelines for AI agents working with this project.

- When using `gh` CLI remember to exit sandbox. Otherwise you will not get access to session tokens.
- When working in a worktree remember to copy over all required config files from the root repo.
- Codex only: after sourcing `.envrc`, add asdf shims to `PATH` so all tools use the versions from `.tool-versions`, for example: `source .envrc && PATH="$HOME/.asdf/shims:$PATH" asdf exec mix test`.

### Always Run Checks Before Submitting Changes

Use the CI configuration file in `{changed_directory}` as a list of checks to run. Prefer variants dedicated for local developments over ci specific ones.

Remember to:

- Never run `*:ci` commands locally. Use variants dedicated to local development.
- Never use `npm clean-install` or `npm ci`. Always use faster variant `npm install` or `npm i`.
- There is no point in running `npm run check-format` and then format again. Just `npm run format` and see what has changed.

### Hard rules

- Never touch any remote environments like production, beta etc without explicit authorization.
- On infrastructure you work strictly read-only: you may run commands that only read state, and any command that changes state you only suggest to the user, who runs it themselves.
- Use `:global.trans` for application-level synchronization. Do not introduce database locks, including PostgreSQL advisory locks or explicit row locks.
- All documentation, comments, tickets etc. have to be in english.
- Follow the existing contract. Do not introduce new defensive conditionals, null checks, or fallback values.
- Comments explain why, not what. Do not write comments that restate what can be read from the code; they get out of sync quickly.
- Review comments are suggestions, not instructions. Verify each one against the code before acting on it and push back on the ones that are wrong instead of implementing them.
- If handling a missing or invalid value is part of the intended behavior, implement it because the contract requires it, not as a guess.
- If you are unsure whether something can be missing, invalid, or optional, inspect the types, schema, and call sites first.
- Enforce invariants and fix incorrect assumptions at the source.
- Legacy code and compatibility fallbacks are not accepted in this codebase. Replace old paths instead of preserving them.
- The frontend uses React Compiler. Do not add `useMemo` or `useCallback` for routine memoization or referential stability; write values and callbacks inline.
- Every commit message must include a meaningful body that explains the change's intent and motivation—not merely what the diff does—so future readers tracing code through `git log` or `git blame` can understand why the change was introduced.
- Never commit without running all CI checks beforehand. Everything must pass, no matter if CI fails because of your changes or not.
- When rebasing newer version prefer origin/ variant and always use git rebase.
- In database migrations, never add defaulted columns or run full-table UPDATEs on existing tables, as these may lock the table for the duration of the migration.
- Tests must not inspect production configuration—including its values, keys, presence, absence, shape, ordering, model lists, thresholds, routing, fallbacks, or selection logic.
- Configuration-only changes must not add or modify tests, except to fix existing tests that violate the rule above. Remove their configuration coupling instead of updating hardcoded expectations.
- When making frontend changes and browser tools are available in your harness you must validate all changes visually by making screenshots on mobile, tablet and desktop resolutions. You must introspect those screenshots and make sure that they look good and do not break any design principles. Attach the screenshots to the PR description.

### Running Elixir Tests Requires .envrc

**Before running any Elixir tests (`mix test`), source the `.envrc` file.**

Use the existing local services and credentials defined in `.envrc` and project config. Do not start replacement services such as Docker Postgres unless they are actually required.

```bash
cd <elixir_project_dir> && source .envrc && mix test
```

JS projects use **npm** as the package manager (indicated by `package-lock.json` file). Install deps outside sandbox to avoid permissions issues.

Import do własnego repozytorium

Strona serwuje wszystkie pliki w surowej postaci, więc agent może je pobrać sam. Najprościej wkleić mu poniższy prompt.

Przez agenta — instrukcja dla agenta leży pod /llms.txt:

Przeczytaj https://TWOJA-DOMENA/llms.txt i zainstaluj w tym repozytorium opisane tam skille oraz AGENTS.md, postępując zgodnie z instrukcją z tego pliku.

Ręcznie, jednym poleceniem — paczka ze wszystkimi sześcioma skillami. Dla Claude Code:

mkdir -p .claude/skills && curl -fsSL https://TWOJA-DOMENA/skills.tar.gz | tar -xz -C .claude/skills

Dla Codex te same katalogi trafiają do .agents/skills/:

mkdir -p .agents/skills && curl -fsSL https://TWOJA-DOMENA/skills.tar.gz | tar -xz -C .agents/skills

AGENTS.md — do katalogu głównego repozytorium. Jeśli masz już własny, połącz je ręcznie, bo to polecenie go nadpisze:

curl -fsSL -o AGENTS.md https://TWOJA-DOMENA/AGENTS.md

Claude Code czyta CLAUDE.md, więc żeby obaj agenci korzystali z jednego zestawu zasad, wystarczy w nim import:

@AGENTS.md

Opcjonalny plik agents/openai.yaml obok SKILL.md ustawia w Codex nazwę wyświetlaną i to, czy agent może sięgnąć po skill sam:

raw
interface:
  display_name: "Babysit and Merge"
  short_description: "Babysit a PR, then squash-merge it and close its ticket"
  default_prompt: "Use $babysit-merge to babysit this PR, then squash-merge it and move its ticket to Done."

policy:
  allow_implicit_invocation: false

Skill wywołuje się po nazwie — /babysit w Claude Code, $babysit w Codex — albo agent sięga po niego sam, gdy zadanie pasuje do opisu w nagłówku.