Vendors are shipping AI-powered investigation assistants, autonomous triage workflows, and natural-language threat hunting interfaces. Much of the market conversation focuses on the AI platform itself, but the single most consistent barrier to adopting agentic workflows is not the AI layer, it is the data layer beneath it.
Without that foundation, AI workflows don’t just underperform, they fail in ways that are difficult to diagnose and costly to fix mid-deployment. Poorly-structured data can force agents into brute-force, trial-and-error querying to compensate, which can multiply token costs well beyond budgeted spend.
Why the Data Layer Decides Everything
Agentic workflows are reasoning engines. They ingest context, form a hypothesis, query for evidence and act; every step is only as good as the data supplied to it. Consider what an AI triage agent needs to investigate a lateral movement alert:
- Normalized authentication events across domain controllers and cloud identity providers
- Process execution telemetry with parent-child relationships intact
- Network flows correlated to endpoint context
- Enrichment data available at query time
If any of those sources are missing, delayed, un-normalized, or siloed, the agent’s output degrades. It has no way to flag what it cannot see, so it simply reasons with what it has.
Detection engineers have understood this problem for years. The agentic SOC era makes it existential, because AI agents amplify the quality of your data in both directions, and they do it with more confidence than a human analyst typically expresses. A well-structured, consistently normalized data layer makes an agentic workflow faster and more accurate than any analyst working the same problem. A fragmented, inconsistently sourced one produces conclusions delivered with that same confidence but are either wrong, because the agent reasoned correctly over incomplete or mismatched data, or outright hallucinated, because the agent filled a gap in the data with something that was never there to begin with. Either failure mode looks identical in the output: a clean, confident disposition that is hard to catch until something breaks.
What Agentic Readiness Requires at the Field Level
“Ready for AI” is not a general statement about log volume or SIEM coverage. It is a specific, testable set of conditions at the field level. In SRA’s view, log sources need to be evaluated against criteria like these before an agentic workflow should be trusted to query them:
- Consistent field naming across sources reporting the same event type, so an agent querying “authentication failures” gets the same schema whether the source is Entra ID, Okta, or an on-premises domain controller.
- Parent-child process lineage preserved end to end, not truncated or dropped during ingestion, so an agent can trace an alert back through the full execution chain without a manual pivot.
- Timestamp normalization to a single time zone and format across every source, since an agent correlating events across systems with inconsistent time handling will misorder or miss the correlation entirely.
- Identity resolution that maps a single human or service account across cloud, on-premises, and SaaS identity providers, so an agent does not treat one person as three unrelated accounts.
- Enrichment data, such as asset criticality, ownership, and threat intelligence context, available at query time rather than requiring a separate lookup the agent has no path to trigger.
- Retention windows long enough and consistent enough across sources that a 30-day correlation query returns 30 days of every relevant source, not 30 days of one and 7 of another.
A source that fails even one of these checks does not disqualify a whole environment, but it defines exactly where a remediation roadmap needs to start; a general statement that an environment “needs better logging” is not specific enough to act on. This is also why agentic SOC readiness is not the same project as general log management maturity: an organization can have excellent SIEM coverage, strong alerting, and a mature detection engineering practice, and still fail on identity resolution alone, because a human analyst can bridge an identity gap mentally in a way an automated agent cannot. The gap only becomes visible once something is asking the data to reason without a person in the loop to catch the mismatch.
The Data Lake-Centric Architecture
Data lake-centric describes a design principle, not a specific topology. The data lake is the center of gravity for how security telemetry is collected, structured, retained, and queried. Every other platform in the stack, including the SIEM, operates as a consumer of that layer rather than a replacement for it. The topology that implements this philosophy will look different depending on the organization. Some environments consolidate into a single lake. Others distribute across multiple lake instances driven by residency requirements, sovereign cloud boundaries, or the cost of moving high-volume telemetry across regions. What stays consistent across all of them is the design principle: the lake owns the data, and everything else queries it.
Distributed ingestion nodes are what make this architecture practical at scale. Nodes sit at the edge, in regional environments, data centers, or cloud partitions, handling collection, parsing, enrichment, and normalization before data moves. Forwarding processed events rather than raw logs reduces egress costs and ties bandwidth to processed event volume rather than raw log volume. For a 15 TB/day environment moving data across cloud regions, the difference between those two approaches can represent a seven-figure annual cost swing. Residency and regulatory constraints are handled at the pipeline level, not the query level: data that cannot leave a boundary never does. The SIEM retains its role in alerting, correlation rule execution, and case management, but in this model, the SIEM becomes a consumer of the lake rather than the source of truth.
Where Ontologies and Semantic Layers Fit
A parallel conversation is happening alongside the agentic SOC pitch: ontologies and semantic layers as the next layer AI agents need to reason reliably. The terms get used loosely, so it’s worth defining them plainly:
- The data layer is the raw material: normalized telemetry like authentication events, process executions, and network flows, sitting in the lake.
- An ontology is a map of relationships between that data, independent of where it came from. It’s what says, “this login and this file execution both belong to the same identity” or “this identity owns this asset,” rather than treating each event as an isolated record.
- A semantic layer takes that map and makes it queryable through one shared, governed set of terms. Instead of a dashboard, a BI tool, and an AI agent each having their own private definition of “authentication failure,” they all ask for the same thing and get the same answer, because the semantic layer is the single translation point everyone queries through.
- The AI layer is the agent itself, and it’s the one that benefits most from this: instead of guessing what a raw field means or reasoning about a relationship that was never made explicit, it queries the governed vocabulary directly and gets a consistent, defined answer back.
Gartner analysts have predicted that 60% of agentic analytics projects relying solely on the Model Context Protocol (MCP) will fail without a consistent semantic layer beneath them, which is really a statement about this same dependency: each layer only works if the one below it is solid.
That is also the limit of what an ontology or semantic layer can do. Both are compiled views of meaning, not a fix for the data underneath. If identity is not resolved across sources, an ontology can label the mismatch consistently, but it cannot merge the accounts. Built on a data layer that already resolves identity, normalizes fields, and aligns retention, though, an ontology and semantic layer are a real accelerant for agentic SOC, worth watching as this space matures.
Federated Search Has a Role, But Not the One Vendors Are Selling
The pitch dominating vendor conversations right now is to leave data where it lives, connect a query layer across your SIEM, EDR, cloud logs, and identity stores, and get a unified view without the cost or complexity of a centralized lake. Federated search looks compelling in a vendor demo. Under production conditions, it does not hold up, for a structural reason rather than an implementation one: every system in a data-in-place model keeps its own retention window, its own API, and its own query language, and a federated layer can only query across them as well as the least capable source allows.
A common example: an organization needs endpoint telemetry for an investigation, only to find the source platform charges extra to expose the logs at all, retains them for a fraction of the window the investigation actually needs, and returns them through a query interface that cannot express the correlation the agent is trying to run. None of that is a bug in the federated layer. It is the federated layer working exactly as designed, which is the point: a query-in-place architecture inherits every retention limit, API constraint, and access restriction of every source it touches, and it has no mechanism to normalize any of that away. The result is an environment where you cannot depend on the data being there when you need it and cannot aggregate or analyze it even when it is.
Federated search does have a legitimate role, but it sits above a well-structured lake, not in place of one. It earns its place querying across multiple lake instances when residency requirements prevent full centralization, and for enrichment sources that don’t need to be primary logging targets: threat intelligence platforms, asset inventories, identity context stores.
The failure mode SRA sees in the market is organizations adopting federated search as a shortcut past the hard work of building the lake. When that happens, the fragmentation, latency, and data quality problems do not go away; every agentic workflow built on top inherits them, the same way a building inherits every shortcut taken in its foundation. Skip the identity resolution work at the start, and an agent six months into production still can’t tell whether a suspicious login and a suspicious file execution belong to the same compromised account, no matter how capable the model behind it becomes. The gap does not surface as a bug to fix. It surfaces as an investigation the agent quietly gets wrong, because the platform can only reason as well as the foundation underneath it lets it.
Two Environments, Two Outcomes
This difference isn’t theoretical: SRA has built data layers across both ends of this spectrum, and the contrast is consistent.
In an environment where every source is queried in place rather than centralized, a typical setup has authentication logs in the SIEM, EDR telemetry in a separate vendor console, network flow data in a third platform, and no consistent identity resolution across any of them. An AI triage agent investigating a lateral movement alert must query three systems with three different schemas and retention windows, and no guarantee that the account name in one system maps to the same person in another. The agent either times out waiting on a federated query, or it returns a disposition built on two of the three data sources because the third was unreachable. The analyst reviewing that output has no way to tell the difference between a genuinely clean alert and one the agent could not fully investigate.
In a data lake-centric environment built along the principles above, the same investigation runs as a single query against normalized, co-located data: authentication history, process execution, and network flows all return in the same schema, on the same timeline, with identity already resolved across sources. An agent working against that environment could complete the same investigation in minutes instead of escalating to an analyst, and the analyst reviewing it could trust that the evidence set was complete. The architecture is the variable. The AI model performing the reasoning is often the same one in both cases. The same foundation pays off well beyond agentic workflows, too: a dashboard or management report becomes a query against clean, unified data instead of a new integration project, rather than requiring an expensive third-party reporting tool.
How SRA Approaches This: The Security Data Pipeline
SRA’s Security Data Pipeline (SDP) service builds the data layer that makes detection engineering, threat hunting, and agentic SOC possible. The methodology is architecture-first: ingestion topology, normalization approach, residency handling, and cost model are all defined before platform selection. That sequencing matters because the foundational design choices made early are what determine whether the data layer holds up once agentic workloads start hitting it at scale, not the platform decisions that are easy to change later. SRA is extending that same architecture-first approach into a dedicated agentic readiness methodology, and the criteria and phases described below reflect where that work is heading.
SRA’s onboarding methodology establishes a structured baseline for every log source: correct parsing, field fidelity, and consistent schema alignment. That baseline is what an agentic readiness evaluation would build on, and it is the direction SRA is actively developing SDP toward: identifying which sources are queryable at the field level an AI agent would require, which have gaps that limit automated reasoning, and what remediation would look like. Cost modeling covers the full stack today: ingestion, storage, pipeline compute, and query, grounded in unit economics benchmarks SRA has developed across client environments, giving clients a realistic projection that includes the operational costs rarely visible in a vendor’s sales process.
What Readiness Actually Looks Like
An AI triage agent receives an alert, queries 30 days of normalized authentication history for the involved account, correlates against endpoint process execution data, joins on network flows, checks threat intelligence enrichment, and produces a disposition recommendation with supporting evidence in under two minutes, without an analyst touching it. That workflow is achievable with the right foundation in place. SRA has clients actively building that foundation today, and each of them started with the data layer rather than the AI platform.
Organizations treating data infrastructure as a security program deliverable today are the ones positioned to operationalize agentic SOC over the next 18 months. Those that skip that step will spend that time troubleshooting why their AI agents keep returning low-confidence results, escalating noise, and missing context that was never in the lake to begin with.
Building the Roadmap: Where to Start
Most organizations do not need to rebuild their entire logging architecture to reach agentic readiness. They need a sequenced roadmap that closes the highest-impact gaps first:
- Baseline every log source against the field-level readiness criteria above, and rank gaps by which sources feed the highest-value agentic use cases, such as identity-based lateral movement detection or endpoint-to-network correlation.
- Remediate the sources with the largest readiness gaps first, prioritizing normalization and identity resolution fixes over new source onboarding, since a poorly normalized existing source degrades every query that touches it.
- Validate readiness with the same queries an agentic workflow would run in production, not with a general health check, so the roadmap is measured against the actual bar the AI platform will need to clear.
This keeps the roadmap grounded in what an agent will query, not a generic data-quality initiative that improves dashboards without improving what an AI workflow can act on.
The Cost of Getting the Sequence Wrong
Organizations that select an agentic SOC platform before addressing the data layer tend to miss the gap during proof-of-concept, not because it isn’t there, but because a small, hand-prepared dataset behaves nothing like the client’s actual production environment at scale. The gap surfaces once the platform hits production: the same normalization and identity resolution issues Security Data Pipeline work uncovers in any data layer assessment show up as a stalled rollout, a platform re-negotiation, or a security leader explaining to the board why an AI investment hasn’t reduced analyst workload the way it’s capable of, because the platform never had the data foundation it needed to do so.
The rework required to fix a data layer after platform selection is almost always more expensive than fixing it first, because the platform’s ingestion assumptions, query patterns, and licensing model are all built around data that turns out not to exist in the form the platform expects. Architecture-first sequencing exists specifically to avoid that outcome.
SRA builds that layer, and SRA’s free SCALR AI is a working example of what an agentic SOC platform looks like when it’s designed around a sound data foundation rather than bolted on top of one. If your organization is evaluating agentic SOC platforms or planning a SIEM modernization that needs to account for AI readiness, contact SRA to understand where your data layer stands and what closing the gaps requires.
SCALR AI is now available for free download and deployment into your organization's Azure environment
Your AI security data stays in your cloud. Don’t take that to mean that SCALR AI only works in a Microsoft environment; it can adapt to work with any of your security tools. Schedule a demo with us to start identifying other opportunities for SCALR AI to enhance your AI security processes.
Jared Anthony
Jared focuses on SIEM and network security toolset implementations and engineering along with emerging services in the cybersecurity industry. He has worked within many industries including pharmaceutical, healthcare, insurance, telecommunications, and retail. He is familiar with leading SIEM platforms including Splunk, QRadar, Exabeam and others.
Jared previously worked in security operations, specifically intrusion detection and incident response. His security operations scope covered telecommunications, pharmaceutical, and financial services industries.
Jared has completed the SANS Intrusion Detection In-Depth course, GIAC Certified Intrusion Analyst (GCIA), GIAC Continuous Monitoring Certification (GMON), and is a Splunk Enterprise Certified Architect.





