Google’s Jules is one of the more interesting tools to come out of the recent wave of developer AI. Unlike in-editor autocomplete or interactive chat panes, Jules runs asynchronously in the cloud: you assign it a task or schedule it, it spins up an isolated VM, clones your GitHub repo, runs the toolchain, and opens a pull request.
It’s designed to be a background worker that takes routine codebase maintenance off your plate. To help users get started, Jules ships with built-in default roles—like Performance, Design, and Security.
And if you inspect how those official roles are written, they follow a very specific template: a 5-step framework packed with exhaustive bulleted checklists. The default Performance role inside Jules is structured roughly like this:
1. PROFILE - Hunt for performance opportunities:
FRONTEND PERFORMANCE:
- Unnecessary re-renders in React/Vue/Angular components
- Missing memoization for expensive computations
- Large bundle sizes (opportunities for code splitting)
- Unoptimized images (missing lazy loading, wrong formats)
- Missing virtualization for long lists
- Synchronous operations blocking the main thread
- Missing debouncing/throttling on frequent events
- Inefficient DOM manipulations
BACKEND PERFORMANCE:
- N+1 query problems in database calls
- Missing database indexes on frequently queried fields
- Expensive operations without caching
- Synchronous operations that could be async
- Missing pagination on large data sets
...
2. SELECT - Choose your daily boost
3. OPTIMIZE - Implement with precision
4. VERIFY - Run format, lint, and tests
5. PRESENT - Open PRWhen you see that the platform’s own creators wrote their default roles this way, the natural conclusion is: "This is the official blueprint. This is how I should structure all my custom roles." It looks thorough. It covers real, tangible engineering problems that humans care about.
The problem is that large language models do not process exhaustive checklists the way human developers do. When you run agents against this pattern across real repositories, two specific failure modes quickly emerge:
1. Selective Blindness (The Anchor Trap)
When you give an LLM a granular checklist of flaws, it stops evaluating the codebase holistically. It shifts from reasoning about system architecture to brute-force pattern matching against those exact bullet points. If a file has an obvious architectural defect, a broken error boundary, or a leaking resource that isn't explicitly mentioned in the checklist, the agent walks right past it. It only searches for what you told it to search for.
2. The "First-Bullet" Hyper-Optimization Trap
Even worse is positional bias. Models naturally overweight the top of lists. If your first bullet points mention "unnecessary re-renders" or "missing memoization," the agent fixates on that specific pattern.
To be fair, Jules actually tracks memory between agent runs specifically to prevent direct cycling. It knows what it changed previously. But because it remains firmly anchored to that initial bullet point, its reaction isn't to look elsewhere—it tries to maximize that single function's performance to the absolute extreme.
"In run one, it wraps the function in useMemo. In run two, it inlines a loop to shave off 2 milliseconds. In run three, it refactors internal variables for another micro-boost. It traps itself in hyper-optimization tunnel vision while genuine debt sits untouched across the rest of the project."
Seeing how the built-in roles behave in practice led me to invert the entire architecture: stop giving the model an anchored list of flaws to hunt, and start giving it a disciplined way to walk the codebase.
That is why I built jules-roles. Instead of packing massive checklists into a few broad roles, the library breaks engineering maintenance down into discrete, specialized personas—like Scribe (documentation), Validator (test suites), Craftsman (cyclomatic complexity), Prism (type integrity), and Sentinel (security auditing).
The architecture rests on three core principles:
1. Zero-Anchor Directives
The prompt defines the discipline and its standards, not a laundry list of specific defects.
- Validator: The instruction isn't "search for these 10 missing test scenarios." The instruction is to inspect the test suite, identify an unasserted boundary or missing integration coverage, and author a complete AAA (Arrange-Act-Assert) test file with proper mocks.
- Craftsman: It doesn't search for specific bad syntax; it measures cognitive branching and flattens nested logic within a single cohesive unit.
Because the model isn't anchored to a prescriptive checklist, it explores the module naturally and addresses the real, unique issues present in that specific code.
2. Persistent Journals (.jules/<role>.md)
To prevent the agent from getting tunnel-visioned on the same function three runs in a row, every role maintains a stateful markdown journal committed directly to the repository under .jules/.
Each run performs an LRU (Least Recently Used) directory traversal:
- It reads its journal to see what files or directories were inspected in previous runs.
- It identifies the oldest, least-recently-visited segment of the codebase.
- It focuses exclusively on that segment.
- It logs its progress back to the journal as part of the PR.
This turns Jules from a random number generator into a systematic background daemon. It steadily traverses every directory in the project instead of clustering around whatever file it looked at yesterday.
3. Strict Atomic Output
Background AI pull requests become noisy and unreviewable when an agent tries to fix types, rewrite tests, and update docs all in a single 40-file diff.
Every role in jules-roles enforces atomic work bound to strict Conventional Commits:
test(auth): add edge case coverage for session expiration
refactor(billing): flatten nested checkout conditionals
fix(api): validate incoming webhook payload signaturesEach run produces one cohesive, isolated change that takes a human maintainer less than two minutes to review and merge.
Checklists are great for human developers because humans have common sense: we use them as reminders while still seeing the bigger picture. LLMs don't. When you give an LLM a checklist, you anchor its attention to a narrow set of keywords and introduce severe positional bias.
If you want an autonomous background agent like Jules to systematically improve a production repository, stop anchoring it with bulleted lists of things to fix. Give it a clear functional discipline, a stateful journal to track its traversal, and the freedom to evaluate code without preconceived biases.
The full repository with all role configurations and traversal specs is open source at github.com/Evgenii-Zinner/jules-roles.