More Than Technology: How Agentic AI Changes Minds, Culture, and the Way We Build
written by Stefan Christoph
- 22 minutes readThis post is the sourced companion to a talk of the same name. The argument: once usable technology is available, the bottleneck often lies in how we think, work together and build.
This Morning, My Agent Had My Day Ready
Most mornings I open one note: my morning brief. My agent prepares it. It lists the day’s meetings with the context I need (who is in the room, what we discussed last time, what is still open), separates the emails that need my answer from those that can wait, and collects news about the customers I meet that day. The agent prepares; it does not send or decide anything, and anything that affects a customer I check myself. Sometimes the brief is wrong; then I correct it and continue with the checked version. I wrote earlier about how I run several agents through the day [9].
Connecting a calendar was the easy part. The harder parts were knowing what a good brief is (nobody could write that down for me), learning which parts to trust, and giving up my old routine. In a team or a company each of these gets harder, because approvals, shared data and shared accountability come on top.
Two September studies from insurance, a regulated industry, show the pattern at organisation scale. Insurance Business reports that KPMG found only 8% of insurers rate their workforce as highly proficient with AI tools, while 54% say they provide effective AI training [1]. Accenture surveyed 263 senior insurance executives: 70% run targeted AI skills initiatives, only 14% have scaled them, and 81% report at least a 5% improvement in gross written premiums while only 23% have achieved enterprise-wide AI integration [2]. The studies use different populations and measures. I read them as an invitation to work on the part that is not technology, on three levels: minds, culture, and the way we build.
Part 1: Minds
People Feel Faster. Companies Don’t, Yet
McKinsey’s State of AI survey from August: 80% of respondents say AI has improved their productivity, but only 37% report a positive contribution to their organisation’s EBIT, and the share of “high performers” (at least 5% of EBIT from AI) stayed flat at about 6% [3]. The first number is about individuals, the others about organisations; together they suggest that individual gains do not turn into company results on their own.
Part of the reason is what changes with agents. In this post, an agent is a system that gets a goal, plans, uses tools, acts and looks at the result, within permissions and human checkpoints. Picture one that reads a change request, drafts code and a test, and opens a pull request; a human decides the merge. McKinsey’s case study on AWS draws the consequence: “When agents become teammates rather than tools, team structures, decision ownership, and performance expectations need to evolve alongside them” [4].
The Skill That Matters Now
For most of our careers the important skill was producing the artefact yourself: the code, the analysis, the test plan. In agent-supported work, four other steps matter more:
- Define what good looks like before you ask. For my morning brief, this was the hardest work.
- Give context: conventions, constraints, the history of a system, everything that is not written down.
- Check what matters, every time.
- Decide. The agent proposes; you decide and own it.

The skill that matters now: before, produce the artefact; now, lead the work
It resembles the work of experienced tech leads and is new for most of us, so train for it deliberately. The 4D framework of AI fluency by Rick Dakan and Joseph Feller is a good aid: Delegation, Description, Discernment and Diligence, with a free Anthropic course for structured practice [5].
Four Fears, Four Answers
In Bitkom’s April survey, 48% of employed people in Germany use AI at work, and 45% fundamentally reject AI support in their job [6]. You can use AI and still have reservations, and scepticism is often a sign that someone takes the risks seriously. McKinsey groups such reservations into four fears [7]. I treat them as rational design inputs, each with its own answer:
- Being found out. People do not know where to start and are afraid to admit it. The answer is making learning safe and normal: a team meeting where someone shows where the agent got it wrong is worth more than a polished demo.
- Losing your professional edge. If an agent does what took me years to learn, what is my contribution? In my experience experienced developers and architects talk about this fear least. The answer is showing what the higher-value work looks like.
- Accountability without control. “If the agent gives me the wrong answer and I act on it, it’s still my mistake” [7]. In a regulated industry this is a design requirement: decision and override points must be visible in the workflow, not only in a policy.
- Stepping into the unknown. The tools change every few weeks. The answer is honest leadership: “I don’t know exactly where we’re going, but I’m learning alongside you” [7].
The article calls what happens otherwise “paralysis dressed up as busyness” [7]. No model update delivers these four answers; they need leaders’ time.
What I Stopped Doing, and What I Started
For a long time my job as a solutions architect was to know things and be well prepared. When my agent took over much of the research and preparation, I asked what was left for me. I stopped preparing customer meetings from scratch, reading every channel myself and hunting for receipts. I now spend more time with customers, on architecture and trade-offs, on writing, on deciding, and on small running prototypes instead of descriptions. My edge did not disappear; it moved towards judging quality, asking better questions and deciding. One rule always applies: everything the agent drafts, I check before it leaves my desk.
The habit that changed most is my first question for every new task: can an agent do this, or help me do it better? The second follows at once: what does good look like, and who checks it?
Inside my own setup, as it is and not as a reference architecture, sits an agent in Kiro CLI [8]. Around it are about 80 skill files (plain-text playbooks the agent loads only when a task matches), about 15 steering files (standing rules, many written after something went wrong) and about 14 MCP servers through which it reaches mail, calendar, team chat, the CRM, documentation and a browser. A notes vault is its memory, the morning brief runs unattended on a cloud development machine before I wake up, and anything that leaves I approve myself. “Give context” has a concrete form here: writing and maintaining skills and steering files is most of the work, more than the programming. The counts are from my files as of October 2026.
In short, Part 1: trust is built, and it has to be maintained. Define what good looks like before you delegate. Name the fears openly; each needs its own answer. Make decision and override points visible in the workflow. Check everything before it leaves your desk.
Part 2: Culture
Where Do You Stand? Most Pillars Aren’t Technology
AWS’s maturity model for adopting generative AI has four levels (Envision, Experiment, Launch, Scale) and six aspects: business, people, governance, platform, security and operations [10]. Only two of the six, platform and security, are technology in the narrow sense, and the model says “each aspect of generative AI adoption should be considered independently” [10]. A team can be at Scale on the platform and at Envision on people. The AWS Cloud Adoption Framework for AI puts it shorter: “Culture is king, even more so when adopting AI” [11]. Where is your organisation on people, and where on the platform?
Even AWS Had to Work at Adoption
The centre of the talk is how AWS rebuilt its own go-to-market organisation with agents, documented by McKinsey in August with Greg Pearson, AWS Vice President of Global Sales [4]. Years of fast growth, with teams empowered to build what they needed, had produced a rich but fragmented landscape: “more than 200 first-party tools, operating across multiple data sources, without a consistent business view,” and opportunity-to-cash processes that “spanned across more than 40 systems” [4].
AWS turned customer obsession inward and treated its own customer- and partner-facing teams as “Customer One” [4]. It did not start with agents. It defined the end state, consolidated hundreds of data sets into 20 foundational ones, and broke down how nearly 40,000 field employees spend their time, role by role [4]. Only then did it build agents into the redesigned workflows, followed by adoption and measurement. AWS estimates that agents have cut customer-strategy preparation by about seven hours [4].

Customer One did not start with agents: agents come fourth
What transfers is the order, not the scale. Regulation and legacy systems can make the data step longer, which is a reason to plan for it, not to skip it.
The wider rollout brought the next lesson. One prospecting capability reached more than 90% adoption among its intended users, but once opened to a much wider audience its overall adoption initially plateaued below 50% [4]. It did not fit every workflow, people did not know which agent to use, and they faced an “onboarding tax” of connecting agents and setting permissions before seeing any value [4]. McKinsey notes this happened “even with strong executive sponsorship, significant technical resources, and a culture built around innovation” [4]. If adoption takes deliberate work even with those advantages, it is something leaders can actively shape instead of waiting for the next model.
AWS answered with people work. Field champions give peer-to-peer support, immersion sessions get leaders using the tools and modelling the change, and AI-powered reporting became part of weekly and monthly business reviews [4]. Organisations that get this right invest up to five times more in change management, and some AWS managers now spend less time reviewing data and more time coaching [4]. My lesson: budget for change, not only for tools.
One Governed Entry Point, With Room at the Edges
Instead of learning a new tool for every task, people need one governed entry point, because an approved path that is also the easiest one gets used. At AWS that entry point is Amazon Quick, which the case study describes as “a unified entry point to agents, knowledge, and workflows” [4], [12]. As a builder I would not draw the lesson “stop building” from 200 tools: good ideas often come from people close to their users. A governed foundation should give teams room to experiment and a clear path to share and scale what works. Part 3 has one public example.
The Governed Path, Technically
For a regulated organisation, “make the governed path the easiest one” needs a technical shape. I see five layers, a pattern before any product:
- Identity: user- or workload-scoped access with narrow permissions, and a traceable answer to on whose behalf an action runs.
- Gateway: tools, including existing MCP servers, sit behind one controlled entry point.
- Policy: every tool call through the gateway is checked outside the agent’s code, so the rules hold even when the prompt changes. A consequential action can require a recorded human approval.
- Observability: traces and telemetry, plus logged policy decisions. Audit evidence still depends on retention, access protection and your controls.
- Evaluation: sampled live interactions scored with built-in or custom evaluators.
The path is agent, gateway, policy, tool, and the platform must prevent unchecked side paths to tools.

A governed agent stack: agent, gateway, policy, tool, no side door
On AWS these layers map to Amazon Bedrock AgentCore, whose services work together or independently [13]. AgentCore Identity implements agent identities as workload identities and lets agents act on behalf of users or by themselves [14], [17]. Policy intercepts traffic through AgentCore Gateway and evaluates each request before tool access, with policies in Cedar or natural language; policies can be session-aware, for example requiring that an approval was granted, and every enforcement decision is logged through Amazon CloudWatch [15]. A new policy can first run in LOG_ONLY and later move to ENFORCE; grant the permission to change that mode only to trusted principals [16]. Around these sit the rest of the platform: Runtime to deploy and scale agents and tools, Memory, Browser, Code Interpreter, Observability with OpenTelemetry-compatible telemetry, Evaluations, Optimization, and, in preview, Agent Registry and Payments [13], [17]. My recommendation is to build the controls as reusable platform components and make the approved path the default for adopting teams.
For Every Euro on Technology: Three on Process, Five on People
McKinsey describes a 1:3:5 pattern in successful AI transformations: for every dollar invested in agentic technology, three on process redesign and five on capability building and adoption. Most companies invert this formula, and “these are not supporting activitiesβthey are the real work” [7]. I use the ratio as a mirror, not a budget formula.

1:3:5 β where does your budget go?
The same article describes five steps from A to E: awareness, belief, commit, develop and enforce [7]. Experience and leaders using the tools on real work come first; metrics, incentives and decision rights come last. Starting with the mandate risks compliance without real use.
One Workflow, and Measure Whether the Work Got Better
The case study recommends focusing on one to three areas and reinventing them end to end, and warns: “AI should not be used to accelerate broken processes” [4]. At AWS that was go-to-market; in IT it could be the path from change request to deployment.
Then measure the right thing. Most agentic transformations, McKinsey writes, “measure log-ins, licenses, prompt volume, and agent invocations. What they don’t measure is whether the work is better” [7]. Ask three questions: was it used meaningfully, did the workflow improve (cycle time, quality, error rate), and did the outcome improve (customer experience, cost per case)? AWS combines workflow telemetry, adoption tracking and A/B testing, and the case study is candid that proving business value is ongoing work [4].
In engineering, use two DORA delivery metrics: change lead time, from commit to production, and change fail rate, the share of deployments that need immediate intervention such as a rollback or hotfix [18]. Add PR cycle time and rework within a predeclared window as diagnostics; merged agent-assisted pull requests show adoption, not success. DORA’s 2025 report describes AI as “an amplifier, magnifying an organization’s existing strengths and weaknesses” [19], which is why these measures matter. Fix the definitions and measure a baseline before the pilot, and ship the telemetry with the capability, not six months later.
Mechanisms, Not Intentions
An Amazon maxim often attributed to Jeff Bezos says that good intentions don’t work, mechanisms do. “We want to use AI” is an intention; a metric reviewed weekly, with an owner, is a mechanism. Andy Jassy wrote in April that you need “the right data, mechanisms, and truth-tellers to deeply question what’s changed” [20]. My request: treat uncomfortable feedback as operational data, not resistance. When someone says the agent does not help, ask why and change the workflow or the context.
In short, Part 2: first the order, then the agents. Fix the order before adding agents. Budget for people, not only tools: 1:3:5. Offer one governed entry point, with room at the edges. Build agents on a governed platform: identity, gateway, policy, observability. Measure the work, not the log-ins.
Part 3: The Way We Build
Feeling Productive Is Not the Same as Being Productive
In early 2025 METR ran a randomised controlled trial with 16 experienced open-source developers on 246 real issues from their own repositories [21]. They expected AI tools to make them 24% faster; they took 19% longer, and afterwards still believed the tools had sped them up [21]. METR says these results no longer reflect current tools, but felt and measured productivity can diverge (more evidence in [22]). Ask whether the feature reaches production faster, in good quality.
Three Experiments at Amazon, One Lesson
Swami Sivasubramanian, AWS Vice President for Agentic AI, published three experiments in June [23]. In a pathfinder project, six senior engineers rebuilt the Amazon Bedrock inference engine, estimated at 30 developers over 12 to 18 months, in 76 days; commits per developer went from 2 a week to 40. In a structured sprint, six engineers on the Prime Video Financial Systems team, with tasks broken down by a senior engineer over three weeks beforehand, produced 556 commits against a baseline of 96 in ten days and cut a 90-week estimate to 24 weeks. And more than 50 Amazon Stores teams worked on their regular backlogs “with no special conditions and no handpicked engineers”: the median productivity gain was 4.5Γ, with some teams above 10Γ in normalised deployment velocity [23].

Three experiments, different measures, first-party results
The lesson: of the 50-plus Stores teams, the 25 that adopted both new tools and new practices outperformed those that only added AI to existing workflows; “the workflow matters, not just the tool” [23]. These are Amazon’s internal results and the three experiments measure different things, so read them side by side rather than adding them up; what counts is the common direction.
Five Practices, and Spec-Driven Development
The article names five practices [23]: build context first (conventions, architecture guidance, repositories an agent can navigate); slow down to speed up, because every high-performing team reported an initial slowdown before compounding acceleration; feed agents a backlog of well-scoped tasks instead of babysitting them; make intent explicit before code is written; and shift testing left, so agents run tests and attempt fixes before code reaches the pipeline. Kiro’s frontier-teams page has a self-check for where a team stands [8].
The fourth practice, day to day, is spec-driven development:
- Requirements: what the result should do and how we recognise “done”. The agent drafts and asks questions; the team decides.
- Design: architecture and interfaces, drafted by the agent and approved by the team.
- Tasks: small, testable steps; the team approves the list.
- Implement: only now does the agent build, task by task.
- Test: the agent runs the tests and attempts bounded fixes; unresolved failures return to the team.
- Review: the agent assists, humans check correctness, security and permissions, interfaces, data changes and operational behaviour, and a named person decides the merge.

Spec-driven development: findings loop back to the spec, and the spec lives next to the code
It is not a waterfall: a failed test or a review finding can change requirements, design or tasks. AWS’s AI-Driven Development Lifecycle (AI-DLC) describes the same rhythm, where AI “creates a plan, asks clarifying questions to seek context, and implements solutions only after receiving human validation”, and stores plans, requirements and design artifacts in the project repository [24]. The spec is versioned context for the next run, and named decision points answer the fear of accountability without control.
How I Review Agent Work
This talk is my example. My agent collected public studies, checked sources and drafted; the angle, structure and decisions were mine. The review has five steps:
- Spec first: a written six-page narrative came before any slide.
- Allowed numbers: every figure is on a list with its source. What is not on the list does not go on a slide.
- Automated checks: number allowlists, links, information leaks and readability. I check each load-bearing claim by hand against the original source.
- A different model: an independent model reviewed three successive versions of the deck and returned 124, 20 and 16 raw findings, neither confirmed-error counts nor a quality score. Among the confirmed errors were an EBITDA figure that came from a different article than the one cited, and a quote mistranslated in the German version.
- I decide, finding by finding: accept, refine or reject. The mechanisms line stays clearly marked as attributed, because there is no primary source.
A different model found issues the drafting pass had missed; that replaces neither source checking nor the human decision.
Three Ways My Agents Failed Quietly
Things still go wrong, and the dangerous failures are quiet. Three I hit myself:
- The silent swap. For videos I use a cloned version of my voice. A fallback path replaced it with a stock voice, and the run reported success. I noticed only when I listened.
- The hidden exit code. Without
pipefail, the pipeline reported only the status of its last command; the command before it had failed, so the failure stayed invisible. - The wrong reviewer. My script started the reviewer from the wrong folder, so the requested agent was not found. The run printed a warning and continued with the default agent, so it succeeded without the required independent review. Current Kiro CLI versions stop with an error in this case.
Two lessons. First, check three things: did the process report success, is the artifact complete and usable, and did it come from the intended model, voice or reviewer? Second, a fallback may lower resolution or speed, but it must never silently substitute the voice, model, author or reviewer; better to fail loudly than to deliver the wrong thing quietly. Both rules are in my steering files today.
A Prototype Created Ownership Before a Roadmap
Werner Vogels, Amazon’s CTO, wrote in June about two-pizza teams [25]. Thomas Delteil, a principal scientist on the Amazon Quick team, spent a night with Kiro building the first prototype of what would become Amazon Quick desktop. At the demo the next day, the first question was “how do I get this on my laptop right now?”, the second “what can I help with?”, and within hours the project had owners for its main parts [25]. When a prototype costs days instead of months, building enough to learn can come first; for production and regulated changes, intent and controls still come first, the harder half I wrote about earlier [26].
Vogels’ closing line: “Two pizzas were always about ownership culture, and the tools have caught up to the culture.” In the same text he warns: “it should be you doing the writing, not your AI” [25].
In short, Part 3: productivity is measured, not felt. Feeling productive is not being productive; measure delivery. The gains come from changing practices, not just from new tools. Spec first: the agent executes, humans decide at defined points. Review what your agents deliver; quiet failures are the real risk. Ownership comes before the roadmap.
Close: The Biggest Lever Is Bringing People Along
You have little influence on the next model and a lot on how your team works, and that is not only true for leaders.
Three questions for your team:
- Which fear is strongest in my team, and what would answer it?
- Where are we on 1:3:5? Is it the other way round, with technology getting most?
- Which of the five practices can we start next week?
One experiment: pick one clearly bounded, approved experiment with an owner and a measure of whether the work gets better. For example, an agent writes draft tests for a single service, one person owns it, and you record a baseline before measuring cycle time, review effort and escaped defects. Then decide honestly: continue, adjust or stop, and share what you learned with a colleague.
One course: for the 4D skills from Part 1, Anthropic’s free course AI Fluency: Framework and foundations [5].
Which of the four fears do you see most in your team, and what has helped you answer it?
This post was drafted with my agent from the talk’s narrative and speaker notes, in the same way as the talk itself; I edited it and every number links to its public source below.
Sources
- [1] Insurance Business β Insurers rate themselves AI leaders, but none has rebuilt distribution around it (KPMG, Sep 2026) β trade-press report on KPMG’s “Unlocking AI value in insurance”: 8% highly proficient workforce, 54% effective training.
- [2] Accenture β How insurers drive revenue by deploying AI with intent (2026) β 263 senior insurance executives: 81% at least 5% GWP improvement, 23% enterprise-wide, 70% skills initiatives, 14% scaled.
- [3] McKinsey β The state of AI in 2026: On the road to ROI (Aug 2026) β 80% personal productivity gains; 37% positive EBIT contribution; high performers flat at about 6%.
- [4] McKinsey β Lessons from our alliances: What AWS’s agentic journey can teach CEOs about rewiring for AI (Aug 2026) β the Customer One case study with Greg Pearson.
- [5] Anthropic Academy β AI Fluency: Framework and foundations β the 4D framework by Rick Dakan and Joseph Feller.
- [6] Bitkom β Ein Drittel nutzt KI mindestens einmal pro Woche (Apr 2026) β 48% of employed people use AI at work; 45% fundamentally reject AI support in their job.
- [7] McKinsey β How to close the agentic adoption gap (Aug 2026) β four fears, 1:3:5, the A-to-E playbook, measuring whether the work is better.
- [8] Kiro β Frontier teams β the self-check for frontier teams; Kiro is also the tool behind my own agent setup.
- [9] Stefan Christoph β From Columbo to Coworker β how I run several agents through the day.
- [10] AWS Prescriptive Guidance β Levels in the generative AI maturity model β four levels, each aspect considered independently.
- [11] AWS Cloud Adoption Framework for AI β People perspective β “Culture is king, even more so when adopting AI.”
- [12] Amazon Quick β product page.
- [13] AWS β What is Amazon Bedrock AgentCore? (Developer Guide) β services that work together or independently.
- [14] AWS β Amazon Bedrock AgentCore Identity (Developer Guide) β agent identities as workload identities.
- [15] AWS β Policy in Amazon Bedrock AgentCore (Developer Guide) β interception at the Gateway, Cedar and natural language, session-aware conditions, CloudWatch logging.
- [16] AWS β Policy enforcement modes (AgentCore Developer Guide) β
LOG_ONLYandENFORCE; who may change the mode. - [17] AWS β Amazon Bedrock AgentCore FAQs β the capability list, including Agent Registry and Payments in preview; Evaluations samples and scores live interactions.
- [18] DORA β DORA’s software delivery performance metrics β five metrics, including change lead time and change fail rate.
- [19] DORA β State of AI-assisted Software Development 2025 β AI as “an amplifier”.
- [20] Andy Jassy β 2025 Letter to Shareholders (Apr 2026) β “the right data, mechanisms, and truth-tellers”.
- [21] METR β Measuring the impact of early-2025 AI on experienced open-source developer productivity (Jul 2025) β 16 developers, 246 issues; expected 24% faster, took 19% longer; results no longer current.
- [22] Stefan Christoph β The Bottleneck Moved: What 10 Studies Say About AI Developer Productivity β the wider evidence on felt versus measured productivity.
- [23] Swami Sivasubramanian β How frontier teams are reinventing AI-native development (AWS Machine Learning Blog, Jun 2026) β the Bedrock, Prime Video and Amazon Stores experiments; five practices.
- [24] Raja SP β AI-Driven Development Life Cycle: Reimagining Software Engineering (AWS DevOps Blog, Jul 2025) β AI plans, asks and executes after human validation; artifacts stored in the repository.
- [25] Werner Vogels β A return to two-pizza culture (All Things Distributed, Jun 2026) β the overnight prototype, ownership, and writing it yourself.
- [26] Stefan Christoph β The Other Half of Two-Pizza β prototyping got easy; production is the harder half.
About the Author
Stefan Christoph is a Principal Solutions Architect at AWS, focused on agentic AI, media & entertainment, and helping builders move from demo to production. He writes about AI architecture, developer productivity, and the future of software.
This is a personal blog. Opinions expressed here are my own and do not represent the views or positions of my employer.
β€οΈ Created with the support of AI (Kiro)