Practice 02 / Agent systems

GÜRAY YILDIRIM / INDEPENDENT PRACTICE

The agent is one component. Operability is the product.

I help companies choose the right workflow, build a bounded pilot and create the evaluation, permission, escalation and cost controls needed to operate it responsibly.

Screenprint collage of task cards moving through tool use, human approval, evaluation, rejection and completion checkpoints
Operating model / 01

Tool use passes through explicit gates. Uncertain work stays observable. Every path has a stop condition.

Before model selection

Should this workflow use an agent at all?

Agents are useful where context, judgment and tool choice vary. Stable rules and predictable paths often belong in deterministic software. The design begins with that distinction.

QuestionWhat it changes
Can success be observed?

If completion cannot be measured, evaluation becomes opinion.

What authority is required?

Permissions should match a bounded task, never the model's theoretical reach.

What can fail safely?

Reversible work can allow more autonomy than financial, legal or production changes.

Where must a person decide?

Human review is part of the workflow, not a fallback added later.

What is a useful completion worth?

Cost per successful task matters more than cost per token.

01

Evaluate

Representative cases, failure taxonomy and task-level success measures.

02

Bound

Least-privilege tools, approval gates, budgets and explicit stop conditions.

03

Observe

Traces, tool calls, latency, cost and enough context for diagnosis.

04

Escalate

Clear paths to a person when confidence, scope or risk crosses a threshold.

Focused engagements

01

Agent Workflow Pilot

For a company with one promising, bounded workflow and a real baseline for the current process.

Outputs

  • Current-process baseline and acceptance criteria
  • Agent-versus-deterministic workflow decisions
  • Working pilot and evaluation set
  • Permission, approval and escalation boundaries
  • Telemetry, latency and cost report
  • Roll out, revise or stop recommendation
02

Agent Reliability & Cost Review

For an existing agent that works sometimes but is expensive, inconsistent or hard to govern.

Outputs

  • Trace, evaluation and failure analysis
  • Context, memory and tool-use review
  • Model and routing recommendations
  • Guardrail and human-escalation design
  • Cost and latency optimization plan
  • Prioritized production changes
03

Agent Operating Controls

For teams that need a reusable way to evaluate, authorize and observe multiple agent workflows.

Outputs

  • Reference control architecture
  • Evaluation and release gates
  • Tool and identity boundaries
  • Trace and audit model
  • Operational runbooks and ownership

Start with the real problem

Start with one workflow and a real baseline.

Bring the current process, its volume, failure cost, data boundary and what a good completion looks like. The first job is deciding whether an agent belongs there.

Discuss an agent workflow