As artificial intelligence continues to integrate into security operations centers (SOCs) for automating workflows, a critical question emerges: does the performance of SOC operations rely more on the large language model (LLM) or the quality of security data it processes? This question is pivotal for security leaders navigating the complexities of modern threat environments.
Research indicates that data quality plays a decisive role in enhancing security operation outcomes. A study published in Frontiers in Artificial Intelligence examined output consistency across various models while the Information Systems journal emphasized the necessity of high-quality data across multiple dimensions, including accuracy and completeness.
An investigation by the Provably Better Data project further corroborates these findings, showing that the quality of data, rather than the choice of model, significantly influences the effectiveness of security operations. Notably, high-fidelity network evidence can enhance security outcomes by two to four times when evaluated across core metrics.
Operational Advantages for CISOs
The implications for Chief Information Security Officers (CISOs) are profound. Better data translates directly into measurable benefits:
- Improved telemetry reduces mean time to respond (MTTR) to threats.
- Effective management of token expenses aids in controlling operational costs.
- Demonstrable ROI in security investments helps reinforce the efficacy of security initiatives within organizations.
Security professionals can leverage these insights to inform architectural decisions regarding SOC investments. For CISOs, justifying enhancements to infrastructure becomes easier, ultimately combating issues like analyst turnover from alert fatigue while providing tangible security metrics to stakeholders.
Testing the Performance Drivers
To dissect the factors influencing AI performance in enterprise security contexts, the Provably Better Data research project employed a controlled testing framework. They established two benchmarks:
- A Capture the Flag (CTF) scenario simulating a Volt Typhoon attack with 44 investigative questions.
- An incident response task grounded in a Salt Typhoon dataset.
To isolate data quality as a variable, the team processed four distinct network telemetry types under uniform conditions:
- Enriched logs from Corelight
- Open-source nDPI firewall logs
- Snort 3 intrusion detection system alerts
- NetFlow connection telemetry
The models—Anthropic Claude Opus 4.6, Google Gemini Pro 3.1 Preview, and earlier versions—were rigorously evaluated for performance consistency with identical prompts across all tests, shedding light on the relationship between data quality and model efficiency.
Understanding the Quality Evidence Barrier
While advanced language models demonstrate strong reasoning capabilities, their effectiveness often remains constrained by data quality. If telemetry lacks critical protocol-level information, AI agents cannot fill in the gaps of what wasn’t captured. The effectiveness of an investigation hinges on the available data.
Consider a specific CTF challenge regarding the NetBIOS computer name associated with IP address 10.110.154.113. The results exemplify the stark differences in performance depending on log context:
- Corelight logs contained the requisite information, yielding the answer FINANCE01 directly from the NTLM log.
- Firewall logs recognized NTLM activity but did not illuminate individual log fields, failing to return the correct answer.
This discrepancy highlights how gaps in telemetry can impede swift operational responses. With incomplete evidence, human analysts often have to validate AI outputs manually, prolonging investigations. Conversely, when comprehensive data is present, AI outputs can lead to immediate actionable insights.
Empirical Results: Data Quality vs. Standard Logs
The empirical data from the research indicates that high-quality logs lead to significantly better investigation outcomes:
In the CTF evaluation, accuracy rates varied widely based on the data source:
- Corelight logs achieved a 95.2% accuracy rate.
- Firewall logs reflected a 58.3% accuracy rate.
- Snort 3 alerts recorded only a 39.4% accuracy rate.
- NetFlow data fell short with a 25.8% accuracy rate.
The impact of richer telemetry on model performance was evident. Corelight enabled LLMs to respond to all 44 questions with direct log evidence, while NetFlow only allowed responses to 15 questions. This led to a CTF score of 4,178.3 points for Corelight compared to a mere 970.0 points for NetFlow.
Similarly, incident response testing showcased significant disparities in overall evidence coverage:
- Corelight logs maintained a 90.3% evidence coverage rate.
- Firewall logs achieved a 61.3% rate.
- NetFlow records lagged behind at 30.9%.
- Snort 3 alerts captured just 21.2% of necessary evidence.
For critical investigative requirements, Corelight logs allowed models to address 91.7% of mandatory queries, whereas NetFlow supported only 18.3%, with Snort logs providing merely 10%.
The conclusion: High-quality data provided a fivefold increase in critical incident visibility over lower-quality records.
Additionally, the efficiency of investigations drastically improved with superior data quality. The model completed a full investigation utilizing Corelight logs in just 14.7 minutes; however, it took 27.0 minutes with NetFlow logs and 26.3 minutes with firewall logs.
The takeaway: Poor quality data can unnecessarily prolong investigations due to repeated retry attempts by the models.
Interestingly, hallucinated outputs—instances where models create plausible but incorrect information—were kept to a minimum when prompts instructed models to mark any missing evidence as unanswerable. Hallucinations occurred only once in Snort 3 data but were non-existent in the Corelight, firewall, and NetFlow datasets. Grounded models that acknowledge evidence gaps keep analysts from pursuing fabricated leads, thus enhancing investigation reliability and speed.
Strategic Recommendations for Security Leaders
Automation powered by AI offers significant efficiencies in threat detection and incident management for enterprise networks. Yet, realize that effective automation is not automatic; it necessitates well-structured, protocol-aware telemetry to be truly effective.
According to the research findings, neither model sophistication nor intricate prompting can compensate for fundamental deficiencies in data quality. Security operations leaders planning future SOC investments should actively prioritize evidence quality over mere model selection if they aim for optimal outcomes.
For further exploration of the detailed methodology and comprehensive metrics, refer to the in-depth analysis available in the Provably Better Data white paper on Corelight’s website.
Corelight Network Detection and Response
Access to superior data can enhance security performance by two to four times. Explore how high-fidelity network evidence is essential for effective AI-driven security. Learn more.