Step-back performance review across a whole project, not just one task
A review that looks across everything a person did across an entire multi-part project, not just one submission, and names the pattern: where they got faster, where they kept stumbling, how they grew.
A performance review that runs at the end of a full multi-part project (made up of several smaller tasks and team conversations), pulling together everything the person did across all of it into one honest read of the whole arc. It names patterns that no single task could show on its own: whether the person got faster over the project, where they kept stumbling on the same kind of issue, how their independence grew, what their writing style says about them.
Feedback on a single task is local, it only talks about that one submission. A whole project is a story across many tasks, and someone needs to step back and tell that story honestly at the end. This review names what the person did well across the whole arc, where they kept getting stuck, and what to push on in the next project, broken down by how points were earned across the different tasks. It's also the point where each AI teammate who worked with the person shares their own perspective separately, instead of flattening everyone's read into one tone.
The system gathers three sources of context at once: prior task-level results, the conversation logs with each AI teammate, and the person's own notes and submissions. These feed multiple AI calls. Each AI teammate's perspective is generated as a separate call so the manager, the senior peer, and any specialists each keep their own voice instead of blurring into one. The main scoring runs through one model with a backup model in case the first is unavailable, and a normalization step corrects the score breakdown if the individual pieces don't add back up to the expected total. Everything is checked against a strict format before it ships out, so nothing malformed ever reaches the person reading it.
Sole author of the real-time review pipeline, the parallel context fan-out, the per-character perspective generators, the output validators, and the score-normalization step.
Want the full technical depth, the tradeoffs, what broke, what I'd do differently? Ask the agent about this project.