Securing GenAI in the Real World: Assessing a Structured GenAI Implementation with Amazon Bedrock

BLOG

Generative AI (GenAI) projects often start small and move fast, with clients moving from initial testing to production and deployment as promising outcomes emerge. Along the way, the business value grows, but so can security risks, especially when sensitive data and high-impact workflows begin to rely on AI model decisions.

TL;DR – Moving Generative AI (GenAI) from experimentation to production requires careful orchestration to ensure it doesn’t expose data or expand attack surfaces.

  • Rapid GenAI and agent development can unintentionally introduce new risks and data exposures.
  • As systems scale, gaps often emerge in access control, visibility and response readiness.
  • A structured evaluation across models, orchestration tools and agent workflows helps organizations move to production with confidence.

A Case Study: Evaluating GenAI in Amazon Bedrock

In a recent engagement, GuidePoint Security completed a security assessment of an enterprise GenAI implementation that used Amazon Bedrock. The team evaluated the client’s architecture against more than 100 security controls. During the review, we also analyzed an AI agent and its orchestration. The client had built the agent with a no-code AI agent builder and implemented it using that tool’s orchestration feature. 

This blog will provide high-level details about the assessment and how it helped identify security risks.

Why a Comprehensive GenAI Assessment Matters

A GenAI environment built on strong cloud foundations can still introduce nuanced security risks based on prompt structures, data handling and agent operation. 

When AI is added to business processes, security is no longer just an IT concern. It is a trust issue. Stakeholders often start with questions such as:

  1. Are we protecting sensitive data?
  2. Can we demonstrate control and oversight?
  3. How prepared are we to respond to unexpected behavior?

These are the right questions, but on their own, they don’t provide the full picture. 

Each spans multiple layers of a GenAI environment, including identity, data flows, model interactions, agent behaviors, logging, infrastructure configuration and incident response processes. 

A comprehensive look at a GenAI environment is important, therefore,  to break high-level questions into structured, measurable checks. This approach helps ensure that important details aren’t overlooked and that the results provide a clear, actionable view of the system’s strengths and weaknesses.

The Assessment Approach

 The client’s goal was to gain clear visibility into their current security posture and understand where to focus next.  The assessment focused on practical decisions such as:

  • Who can access the application?
  • What data can it touch? 
  • How is their data protected from the identity and network perspectives?
  • How are agent-based workflow activities and orchestration monitored? 

During this process, GuidePoint created a threat model exercise to determine high-priority security risks. From there, the client gained a clear understanding of what to fix first. 

What Does Threat Modeling for a GenAI Application Look Like?

Below is a sample screenshot of what a threat model diagram looks like using a fictitious AI Chatbot application in an AWS cloud environment. Note that these concepts translate to any GenAI application and are applicable across all cloud platforms.

Image: Example threat modeling architecture in a GenAI assessment for an Amazon Bedrock implementation

This assessment reviewed the client’s application and mapped security controls to all applicable resources deployed for the application. It kept the conversation grounded in business outcomes. The team focused on the technical remediation details to ensure the client’s developers and security teams would have a clear and actionable roadmap.

Which Critical Controls Are Evaluated in a GenAI Assessment?

In the early stages of GenAI experimentation, organizations often grant broad permissions so they can move quickly. Unfortunately, those permissions often remain even after an application or system becomes business-critical.

Access control implementation as part of the assessment ensures strong identity security prior to production launch. This part of the assessment involved:

  • Addressing agent entitlements (i.e., agents access to resources)
  • Ensuring access controls were aligned to job roles
  • Validating end-to-end traceability of user activity, data usage and other evidence

From there, the team focused on data protection, including the information that moves through prompts, responses and logs. Many organizations do a solid job safeguarding databases and storage, but GenAI introduces new places where sensitive data can appear. For example, because prompts and responses can potentially handle sensitive data, they should be treated just like any other data source. Therefore, the controls assessment looked for:

  • Clear rules for what information should and should not be included
  • Practical guardrails around retention (including saved troubleshooting and incident response datapoints)

The evaluation also addressed critical control domains, including (but not limited to):

  • Infrastructure security risks
  • Data poisoning risks
  • Anomaly detection (e.g., malicious agent behaviors)
  • Network security risk

What is Different About Response Readiness in GenAI Environments?

A key part of a GenAI assessment revolves around visibility and response readiness. GenAI systems can behave in unexpected ways. The organization’s preparedness can make the difference between quick incident reconstruction and response and the alternatives, e.g., data loss, exposure, model poisoning or breach. 

In the client environment assessment, GuidePoint determined whether the right signals were being captured and if the teams could investigate incidents without guesswork. This structured approach kept the assessment direct and decision-focused. It also made it easier for the client to prioritize fixes based on severity and impact to its business. They could then move forward with confidence rather than getting stuck in a long list of technical observations.

GuidePoint also conducted a series of application security assessments against the client's scoped GenAI-powered applications. Testing included checks for frequently identified, AI-specific risks such as prompt injection, insecure output handling and sensitive data exposure through model responses.  The output of this was a blueprint of how the applications functioned which allowed for running tests where it mattered the most.

How Are AI Agents Included in the GenAI Assessment?

The most important shift we saw came from agent-driven workflows. When orchestration tools are introduced, the system is no longer just generating text. Agent orchestration allows the system to now take steps, call tools and trigger actions. In the workload reviewed, the key question was whether actions taken by the agents were constrained. We looked for clear boundaries on what actions an agent could take and how approvals were handled for higher-impact actions.

If done properly, agents can accelerate and make work more efficient while staying within guardrails. Otherwise, they can become an automation layer that is difficult to audit and hard to immediately stop when something unexpected happens. 

Our general recommendation is not to treat applications leveraging GenAI models as normal applications, as this could lead to missing where the real risks hide (e.g., prompts, retrieved context, agent actions, etc.)

Move Forward with Confidence

A structured GenAI assessment against a broad control set, combined with an application security review, can help security leaders gain rapid visibility and control over evolving AI environments. They help stakeholders understand:

  • What is safe in the early, experimental stages of GenAI and agent-based development 
  • Security gaps that automated tools might miss
  • What needs to be fixed before going live
  • How to prioritize remediation to shrink the attack surface and protect data without slowing innovation

It can help security, business and IT leadership move forward, building AI solutions into their environments with greater confidence.

When scaling GenAI and agent-based implementations beyond a pilot or test phase, contact us to validate that your workloads and agent orchestration layers are secure, governed and ready to scale and reduce security risks before it becomes a business setback.

Senior Cloud Security Consultant
GuidePoint Security

Mukhtar Kabir is a Senior Cloud Security Consultant at GuidePoint Security. He holds a Bachelor of Science degree in Computer Networks and Security from the University of Maryland Global Campus. He is a Certified Information Systems Security Professional (CISSP) and a Certified Cloud Security Professional (CCSP). Here at GuidePoint, he helps organizations address their cloud security challenges by designing and recommending secure, scalable solutions that adhere to industry security best practices. Mukhtar is also AWS Solutions Architect Professional and AWS Security Specialty certified. He occasionally speaks at major industry conferences, sharing insights on cloud security and best practices. In his words, “I enjoy making complex security concepts easy to understand.