Beyond Static Rules: Building Agentic Content Evaluation Systems on Amazon BedrockHow to turn scattered guidelines into adaptive workflows that score, explain, and improve content at scale

Beyond Static Rules: Building Agentic Content Evaluation Systems on Amazon Bedrock

Everyone talks about generating content with GenAI.

Almost no one talks about how to evaluate it.

And that’s where most systems break.

Generative AI has dramatically accelerated how quickly teams can create content. But in doing so, it has quietly introduced a harder problem:

How do you evaluate that content consistently, repeatedly, and at scale?

This is the part that often gets underestimated. Generating content is the easy demo. Governing its quality is the real engineering challenge.

Across enterprises, this challenge shows up in different ways. Marketing teams need content to follow tone and messaging guidelines. Compliance teams must ensure documents include required language. Platform teams want internal documentation reviewed for quality and consistency. Product teams care about SEO and accessibility. ML teams need dataset artifacts validated for completeness and policy alignment (See Table 1 for more examples).

The pattern is always the same: there is a standard — but it’s fragmented. It lives across documents, tribal knowledge, examples, policies, and reviewer habits. Manual review doesn’t scale. Static rules are too brittle. One-shot prompting is flexible, but often inconsistent and hard to reproduce.

What’s missing is not another prompt — it’s a system.

A system that can define how evaluation should work, refine that definition over time, and apply it consistently.

In this post, I explore a workflow that does exactly that: it generates a structured evaluation guideline (in a machine-readable format such as JSON — though it could also be represented in Markdown or similar formats), refines it iteratively using new inputs, applies it consistently to content, and validates the results using LLM-as-a-Judge techniques.

Along the way, it introduces key capabilities such as drift detection, external enrichment, benchmarking, and full traceability.

The result is a shift away from “ask the model to review this” toward something far more robust:

A system that can build its own evaluation standard, improve it over time, and check whether its own judgments remain trustworthy.

That’s where this approach becomes interesting — not just as a solution, but as a pattern for building more reliable, agentic GenAI systems.

Table 1: Pottential Applications of This Pattern

Why this problem matters more in the GenAI era#

Before LLMs, content evaluation was typically handled in one of three ways.

The first was deterministic validation: schemas, regexes, linting rules, checklists, and business logic. These approaches work extremely well when the problem is narrow and clearly defined. However, they start to break down when evaluation involves softer dimensions such as tone, ambiguity, clarity, structure, or domain-specific nuance.

The second approach was human review. Humans are flexible and context-aware, making them ideal evaluators. But they are also expensive, slow, and inherently inconsistent across reviewers, teams, and time.

The third, more recent shortcut is to rely on a single LLM prompt — essentially asking: “Review this content and tell me if it’s good.” This can work surprisingly well in prototypes, but it often leads to unstable and hard-to-reproduce results. The evaluation criteria are usually implicit, buried in prompt wording, and rarely versioned or traceable.

What teams actually need is not just an LLM reviewer, but a repeatable and structured evaluation system.

That’s why an architecture that addresses these limitations is so compelling. Instead of relying on ad hoc evaluations, it breaks the problem into clear stages: define the evaluation standard, refine it incrementally, apply it consistently, and validate the results. In practice, this approach is much closer to how mature, production-grade GenAI systems should be designed.

Reference Architecture for Agentic Content Evaluation
Figure 1: Reference Architecture for Agentic Content Evaluation

A concrete example#

To make this more tangible, let’s walk through a simple scenario:

Imagine an e-commerce company that needs to review 10,000 product descriptions every day.

Each description is supposed to follow internal guidelines:

  • include key product attributes,
  • maintain a consistent tone,
  • avoid restricted or misleading terms,
  • and meet marketplace compliance requirements.

Today, this process is handled through a mix of manual review, scattered checklists, and a few brittle rules.

Some descriptions are too vague. Others are missing critical information. Some violate tone or compliance guidelines. The result is inconsistent quality — and a process that doesn’t scale.

Now imagine replacing that with a system that can:

  • define what “good” looks like,
  • apply that definition consistently across all content,
  • and continuously improve its own evaluation criteria over time.

That’s exactly the kind of system we’ll build throughout this post.

The agentic idea hiding inside this architecture#

At first glance, this might look like a content-quality pipeline. But there is a deeper GenAI pattern behind it.

This design is really an example of an agentic evaluation loop.

The system is not merely generating an answer. It is interacting with a sequence of artifacts and decisions:

  • it creates a structured representation of the rules,
  • it revises that representation when new evidence appears,
  • it applies the rules to a target artifact,
  • it critiques the result,
  • and it can compare itself to ground truth.

That is much closer to an agent workflow than to a plain prompt-response interaction.

The source document lays out this pattern explicitly: the workflow begins with source files, extracts a guideline, refines that guideline iteratively, evaluates an input content piece, and can then use an LLM-as-a-Judge plus ground truth for validation. It also calls out optional enrichment by an external agent and traceability of the sources used across steps.

In GenAI terms, several important patterns show up here.

1. Structured output as a control surface#

The first pattern is structured output. Instead of leaving the evaluation standard buried inside a prompt, the system materializes it as JSON.

That matters because JSON is not just an output format. It becomes a control surface for the whole system:

  • it can be versioned,
  • diffed,
  • audited,
  • reused across runs,
  • enriched with metadata,
  • and compared across iterations.

When GenAI systems become operationally important, structured artifacts like this are often the difference between “clever demo” and “maintainable system.”

2. Evaluator-optimizer loops#

A second pattern is the evaluator-optimizer loop. One model pass creates or refines a guideline, and a later pass evaluates content against it. Sometimes a further pass judges the quality of that evaluation.

This is more reliable than a single monolithic prompt because the system separates responsibilities. One step focuses on defining standards. Another focuses on applying them. Another focuses on judging whether the application was sound.

That decomposition reduces prompt overload and gives you more places to measure quality.

3. LLM-as-a-Judge, but grounded#

The document explicitly includes an LLM-as-a-Judge stage and recommends comparing outputs against expected or ground-truth references where available.

That is an important realism check. LLM-as-a-Judge is powerful, but it should not be treated as magic. It works best when grounded by one or more of these:

  • a clearly defined rubric,
  • reference outputs,
  • comparison baselines,
  • benchmark datasets,
  • human spot checks.

Used well, it becomes a scaling tool for evaluation. Used casually, it becomes another layer of unverified opinion.

4. Tool-using agents instead of pure prompting#

The workflow also maps naturally to a tool-using agent pattern. An agent can retrieve source files, read prior guideline versions, call the model for refinement, persist artifacts, compute diffs, compare scores, and trigger a judge step.

This is where modern agent frameworks become useful. AWS describes Strands Agents as a model-driven SDK where developers define prompts and tools, and the model plans and executes the next steps using those tools. AWS also describes Amazon Bedrock AgentCore as a platform for building, deploying, and operating agents securely at scale, with services such as runtime, memory, gateway, browser, and code interpreter.

That combination fits this content-evaluation pattern well: Strands can express the agent logic and tool use, while AgentCore can help operationalize it.

A better mental model: don’t ask the model to evaluate content#

A useful reframing is this: Do not ask the model to directly decide whether content is good. Ask the model to first define how goodness should be measured, then make it follow that structure.

Key Insight#

Don’t ask the model to evaluate content. Ask it to define how evaluation should work — and then follow that structure.

That is a much stronger pattern. The moment the standard becomes explicit, you gain several advantages:

  • consistency improves,
  • outputs become easier to compare,
  • errors become easier to debug,
  • version drift becomes visible,
  • and benchmarking becomes possible.

This is also one of the most important “GenAI reality” lessons. Reliability usually does not come from a better prompt alone. It comes from introducing more structure around the model.

A practical agentic architecture#

A realistic implementation on AWS could look like this:

1. Data Storage (Amazon S3)
Amazon S3 stores all core artifacts, including source files, guideline versions, evaluation outputs, and benchmarking results. This ensures durability, traceability, and easy version management.

2. Model Layer (Amazon Bedrock)
Amazon Bedrock provides the foundation models used across the workflow — for guideline generation, content evaluation, and LLM-as-a-Judge validation.

3. Orchestration Layer (Strands Agent)
A Strands-based agent orchestrates the end-to-end process. It leverages tools to:

  • Read and process source files
  • Generate and persist structured JSON guidelines
  • Compare revisions across iterations
  • Compute evaluation metrics and scores

4. Agent Runtime & Production Layer (Amazon Bedrock AgentCore)
Amazon Bedrock AgentCore provides the production-grade environment to run the agent reliably at scale, including:

  • Session management and memory
  • Secure tool access
  • Scalable execution
  • Advanced capabilities such as code execution or browser-based enrichment

5. Compute & Deployment Options (AWS Lambda / Amazon ECS)
Supporting components can be deployed using AWS Lambda for event-driven workloads or Amazon ECS (with Fargate) for containerized, long-running processes, depending on system requirements.

The most important thing here is not a specific service choice. It is the separation of concerns:

  • artifact generation,
  • artifact refinement,
  • artifact application,
  • artifact validation.

That separation is what makes the workflow stable enough to evolve.

Example 1: a simple Strands agent for guideline extraction#

In our running example, the system begins by analyzing existing product guidelines, past high-quality descriptions, and relevant compliance rules. From these inputs, it generates a structured guideline — represented as JSON — that defines what a “good” product description looks like, including required attributes, tone, and constraints.

The goal of this first step is not to evaluate content yet, but to explicitly define the evaluation standard.

The following example shows the general pattern. Strands allows you to define agents with tools, and its Python SDK makes it easy to create custom tools using the @tool decorator. This enables the agent to interact with external data (e.g., reading source documents or persisting outputs) as part of its reasoning process.

from strands import Agent, tool
import json

@tool
def load_source_text(path: str) -> str:
    """Load a source guideline document from local storage or S3-mounted path."""
    with open(path, "r", encoding="utf-8") as f:
        return f.read()

@tool
def save_guideline(path: str, guideline_json: str) -> str:
    """Persist a guideline JSON artifact."""
    with open(path, "w", encoding="utf-8") as f:
        f.write(guideline_json)
    return f"Saved guideline to {path}"

SYSTEM_PROMPT = """
You are a content-evaluation architect.

Your job is to extract a domain-specific evaluation guideline and return it as strict JSON.
The JSON must include:
- domain
- version
- categories
- checks
- scoring_approach
- rationale_summary
- source_references

Do not return prose outside the JSON.
"""

agent = Agent(
    system_prompt=SYSTEM_PROMPT,
    tools=[load_source_text, save_guideline]
)

result = agent("""
Read the source file at ./inputs/seo_guidelines.md.
Create a normalized evaluation guideline for HTML pages.
Then save it to ./artifacts/guideline_v1.json.
""")

print(result)

This example is intentionally simple, but it highlights a critical design choice:

The model is not asked for a vague review. It is asked to produce a durable, structured evaluation artifact.

That shift — from implicit judgment to explicit standards — is what enables consistency, traceability, and continuous improvement in later stages of the system.

Example 2: iterative refinement as a first-class pattern#

One of the most important ideas in this approach is that guideline generation should be iterative, not fully regenerated on every run. This significantly reduces cost, preserves prior knowledge, and allows the evaluation standard to evolve in a controlled and transparent way.

In our running example, new information is constantly introduced — such as updated compliance rules or newly curated high-quality product descriptions. Instead of rebuilding the guideline from scratch, the system incrementally refines it, improving its understanding of what “good” looks like over time.

This turns the guideline into a living artifact, rather than a disposable output.

Strands supports this pattern through hooks and re-invocation mechanisms, which make it straightforward to implement refinement loops. These capabilities align naturally with the evaluator–optimizer pattern, where each iteration builds on the previous one.

from strands import Agent
from strands.hooks import AfterInvocationEvent

MAX_REFINEMENT_ROUNDS = 3
iteration = 0

agent = Agent(system_prompt="""
You maintain a content evaluation guideline in JSON.
At each step, refine the current guideline using the newly provided source text.
Preserve previous valid categories unless the new source clearly contradicts them.
Always return valid JSON only.
""")

async def refinement_hook(event: AfterInvocationEvent):
    global iteration
    if iteration < MAX_REFINEMENT_ROUNDS and event.result:
        iteration += 1
        event.resume = (
            f"Review your previous JSON and improve it using the next source file. "
            f"This is refinement iteration {iteration} of {MAX_REFINEMENT_ROUNDS}. "
            f"Keep the output strict JSON."
        )

agent.add_hook(refinement_hook)

result = agent("""
Start from guideline_v1.json and refine it with the policies in ./inputs/accessibility_rules.md.
""")

Conceptually, this is the shift from a static reviewer to an adaptive system.
The guideline is no longer a one-time artifact — it evolves with new data, new constraints, and better examples, becoming progressively more accurate and useful over time.

Example 3: using AgentCore Code Interpreter to verify scoring logic#

Once the system begins producing structured evaluation reports, an important question emerges:

How do we verify that the scoring logic is actually correct?

This is where deterministic validation becomes valuable.

In our running example, the system may evaluate thousands of product descriptions and generate JSON reports with section-level scores, aggregate scores, and drift indicators across versions. While an LLM can reason about the results, some parts of the workflow — such as score calculations, deltas, and metric comparisons — are better handled with code.

This is an important production pattern: use the model for judgment and interpretation, but rely on tools for deterministic verification.

With AgentCore Code Interpreter, the agent can execute Python code in a managed environment to validate calculations, compare outputs across versions, and identify significant changes in the results. This helps ensure that the evaluation pipeline is not only intelligent, but also auditable and mathematically consistent.

from strands import Agent
from strands_tools.code_interpreter import AgentCoreCodeInterpreter

code_interpreter_tool = AgentCoreCodeInterpreter(region="us-west-2")

SYSTEM_PROMPT = """
You are an evaluation auditor.
When given evaluation reports in JSON, use Python code to:
1. verify score calculations,
2. compare results across versions,
3. compute drift deltas,
4. summarize statistically significant changes.

Return a concise explanation plus the validated metrics.
"""

agent = Agent(
    tools=[code_interpreter_tool.code_interpreter],
    system_prompt=SYSTEM_PROMPT
)

response = agent("""
Load two JSON reports:
- ./artifacts/eval_report_v3.json
- ./artifacts/eval_report_v4.json

Verify whether the overall score is correctly derived from section scores.
Then compute the delta by category and identify the three largest changes.
""")

print(response.message["content"][0]["text"])

This example highlights an important principle for agentic systems in production:

Let the model reason about what should be inspected, but let tools verify what can be computed exactly.

That combination makes the system more robust, transparent, and trustworthy — especially when evaluation outputs begin to influence business decisions or downstream automation.

Example 4: operationalizing the agent with Amazon Bedrock AgentCore#

So far, we’ve focused on building the workflow. But in practice, the hardest part is not creating the system locally — it’s running it reliably in production.

Evaluation is rarely a one-off task. In real-world scenarios, the system must handle repeated workloads, maintain context across sessions, ensure traceability, and operate securely at scale.

This is where Amazon Bedrock AgentCore becomes relevant.

AgentCore provides a framework-agnostic environment for deploying and operating agents in production. It offers capabilities such as:

  • managed runtime and scaling
  • session handling and memory
  • controlled tool access
  • secure execution environments
  • support for advanced capabilities like code execution or browser-based enrichment

Together, these features allow you to move from a prototype to a production-grade agentic system.

A simplified invocation example using the invoke_agent_runtime API looks like this:

import boto3
import json

client = boto3.client("bedrock-agentcore")

payload = json.dumps({
    "prompt": """
    Evaluate ./inputs/page.html against guideline_v4.json.
    Return JSON with section scores, mismatches, suggestions, and rationale_summary.
    """
}).encode()

response = client.invoke_agent_runtime(
    agentRuntimeArn="arn:aws:bedrock-agentcore:us-west-2:[id :) ]:runtime/my-evaluator",
    runtimeSessionId="content-eval-session-001",
    payload=payload
)

print(response)

In our running example, this could represent a production system evaluating thousands of product descriptions per day — each request executed as part of a managed session, with full traceability and controlled access to tools and data.

This matters because evaluation is not just a technical step — it’s an operational capability.
You need to support:

  • repeatable and scalable execution
  • consistent environments across runs
  • auditability of results
  • safe rollout of changes to guidelines and evaluation logic

Moving from local workflows to AgentCore is what transforms this pattern from a clever prototype into a reliable, enterprise-ready system.

Why this pattern is especially relevant now#

The use cases listed in Table 1 are broad for a reason. Whether it’s code review, policy compliance, SEO and accessibility audits, academic writing, brand guideline enforcement, e-commerce validation, incident reports, or dataset quality checks — they all share the same underlying structure.

Across these domains, the problem consistently involves:

  • a changing standard,
  • imperfect and evolving source material,
  • a mix of objective and subjective criteria,
  • and the need for scalable, repeatable evaluation.

These are inherently difficult problems to solve with traditional approaches.

This is precisely where GenAI can provide real value — but only when it is wrapped in structure. Without that structure, you get inconsistent and hard-to-trust outputs. With it, you unlock systems that are not only flexible, but also reliable, auditable, and capable of improving over time.

The production lesson: trust comes from process, not from eloquence#

A final point is worth making explicitly.

In early GenAI prototypes, teams often get impressed by eloquent model output. In production, eloquence matters much less than process.

A trustworthy evaluator is not one that sounds smart. It is one that:

  • uses explicit rubrics,
  • stores evaluation artifacts,
  • can be benchmarked,
  • shows its evidence,
  • supports drift analysis,
  • and can be cross-checked by a second mechanism.

That is why the original proposal’s emphasis on ground truth, benchmarking, traceability, and version drift is so strong. Those are not “nice to have” details. They are what makes the architecture credible.

In this scenario, what started as a manual review process becomes a scalable, self-improving system — capable of evaluating thousands of product descriptions daily while continuously refining its own standards.

Conclusion#

The real opportunity here is bigger than content QA.

This pattern points toward a broader class of GenAI systems where the model does not just generate outputs. It helps define standards, improve them, apply them consistently, and measure how reliable the process remains over time.

That is a far more useful direction for enterprise GenAI.

By combining structured JSON guidelines, iterative refinement, tool-using agents, LLM-as-a-Judge validation, and production infrastructure such as Amazon Bedrock AgentCore, teams can move from one-off review prompts to evaluation systems that are adaptive, auditable, and operationally realistic.

And in my opinion, that is where a lot of the real value of agentic AI will be found over the next few years: not only in generating more content, but in building systems that can govern quality at the speed of generation.

Originally published on Loka Engineering on Medium.

Tags

Topics