Splunk Observability Cloud

Missing events

Wanpa
New Member

We’ve confirmed that the Funding Auth Service successfully processed the transaction, the database status is ACQUIRED, and CloudWatch contains the expected logs. However, some of those log events are missing in Splunk. I’d like to understand where the log pipeline is breaking. It is anticipated that the root cause could be updating new splunk deployment server name with different port and updating splunk ip CIDRs on the terraform prod files.

Labels (1)
0 Karma

ITWhisperer
SplunkTrust
SplunkTrust

Can you determine which logs are "missing"?

Are they from the same time period?

Are they the same format?

Are they the same event types?

etc.

You could do a bit more analysis to give us some clues as to where your issue(s) might be.

0 Karma

Wanpa
New Member

We are investigating an intermittent production log ingestion issue in our AWS ECS environment. The application processes transactions successfully—the order status is updated to ACQUIRED, the outbox event is generated and delivered, and no application exceptions are observed. The expected application log message, “Funding details marked as ACQUIRED and outbox event generated,” is present in AWS CloudWatch but is missing in Splunk for some FOIDs (for example, 00118def-0af3-a6b8-9188-c1725237e3fb). Other log events for the same transaction are searchable in Splunk, making the issue intermittent rather than a complete ingestion failure. We have also verified that similar transactions are successfully indexed in Splunk, while only specific FOIDs are affected. We recently updated the Splunk deployment server configuration and Splunk CIDR/IP configuration, although we have not yet established whether those changes are related to the issue. Based on this behavior, could you help identify where the most likely point of failure is in the CloudWatch-to-Splunk ingestion pipeline? Specifically, would you recommend focusing on forwarder/HEC connectivity, inputs.conf/outputs.conf, deployment server configuration propagation, host-specific collector issues, indexing/parsing, or another component? Any guidance on additional evidence or diagnostics that would help isolate the root cause would be greatly appreciated

0 Karma

Wanpa
New Member

Here’s a professional paragraph you can post on the Splunk Community or send to a Splunk engineer:

 

We are investigating an intermittent production log ingestion issue in our AWS ECS environment. The application processes transactions successfully—the order status is updated to ACQUIRED, the outbox event is generated and delivered, and no application exceptions are observed. The expected application log message, “Funding details marked as ACQUIRED and outbox event generated,” is present in AWS CloudWatch but is missing in Splunk for some FOIDs (for example, 00118def-0af3-a6b8-9188-c1725237e3fb). Other log events for the same transaction are searchable in Splunk, making the issue intermittent rather than a complete ingestion failure. We have also verified that similar transactions are successfully indexed in Splunk, while only specific FOIDs are affected. We recently updated the Splunk deployment server configuration and Splunk CIDR/IP configuration, although we have not yet established whether those changes are related to the issue. Based on this behavior, could you help identify where the most likely point of failure is in the CloudWatch-to-Splunk ingestion pipeline? Specifically, would you recommend focusing on forwarder/HEC connectivity, inputs.conf/outputs.conf, deployment server configuration propagation, host-specific collector issues, indexing/parsing, or another component? Any guidance on additional evidence or diagnostics that would help isolate the root cause would be greatly appreciated.

0 Karma

PickleRick
SplunkTrust
SplunkTrust

OK.

1. Do not just blindly copy-paste some AI slop. And please use some formating, break the text into paragraphs - generally, make the thing readable.

2. You posted in Observablility Cloud section of the forum but your AI-generated content is pointing to a Splunk Enterprise/Splunk Cloud specific features. So make up your mind - which is it?

0 Karma

Wanpa
New Member

The application processes transactions successfully. For the affected transactions, the database is updated correctly (status becomes ACQUIRED), the outbox event is generated and delivered successfully, and the expected application log is visible in AWS CloudWatch. However, for a subset of Funding Order IDs (FOIDs), the same log event is not searchable in Splunk. This suggests the application generated the log successfully and CloudWatch collected it, but the event is either not reaching Splunk consistently or is not being indexed/searchable.

 

During our investigation, we attempted to reproduce the issue in a performance environment before deploying a fix. We increased the logging buffer size from 1 MB to 8 MB because we suspected that high TPS might be contributing to log forwarding delays or dropped events. We then executed load tests in the 500–600 TPS range. Surprisingly, we observed approximately 1,000 records in the outbox database table, while the corresponding Splunk search returned more than 100,000 events. This is the opposite of what we expected—we were anticipating fewer events in Splunk if logs were being dropped. We are now validating whether we are comparing distinct FOIDs versus total log events, and we have modified the load test to generate unique FOIDs to eliminate duplicates.

 

We are also testing with gradually increasing TPS (25, 50, 75, 100, etc.) to determine whether there is a threshold where database counts and Splunk counts begin to diverge. At the infrastructure level, we recently made changes related to Splunk Deployment Server configuration, endpoint/CIDR updates, and Terraform-managed infrastructure, although we have not established any direct correlation between those changes and the observed behavior.

 

Given these observations, how would you approach isolating the root cause? Would you first investigate duplicate ingestion (multiple forwarders or monitor stanzas), forwarder buffering, inputs.conf/outputs.conf, deployment server bundle consistency, HEC/Universal Forwarder behavior, indexing/parsing, or query methodology? Are there specific Splunk internal logs, metrics, or diagnostics you would recommend reviewing to determine whether the issue is due to ingestion, indexing, duplicate forwarding, or simply how the events are being counted?

 

Any guidance or similar experiences would be greatly appreciated. This is a production observability issue, and we’re trying to narrow the exact failure point before concluding the RCA.

0 Karma
Got questions? Get answers!

Join the Splunk Community Slack to learn, troubleshoot, and make connections with fellow Splunk practitioners in real time!

Meet up IRL or virtually!

Join Splunk User Groups to connect and learn in-person by region or remotely by topic or industry.

Get Updates on the Splunk Community!

Guided Onboarding with Auto-schema Is Now Generally Available

  We are excited to announce the General Availability of Guided Onboarding with Auto-Schematization ...

ATTENTION: We’re Moving! (AGAIN!)

The Splunk Community Slack is undergoing a system migration to keep our workspace secure and ...

Deep Dive: Optimizing Telemetry Pipelines in Splunk Observability Cloud

In this session, we will peel back the layers of Splunk Observability Cloud’s cost-optimization features. ...