Getting log formats right is the whole job
A detection rule written against a format that does not match the real device passes in the lab and never fires in production. Notes from building 16 log sources: CEF that is not CEF, positional CSV where order is everything, and a published sample that turned out to be hand edited.
A rule written against the wrong format is worse than no rule
When I started building LogGen, I assumed the hard part would be the Go. It was not. Writing the code for a log source takes an afternoon. Establishing what the device actually puts on the wire takes days, and that is the part that decides whether the tool is useful or actively harmful.
The failure mode is quiet. You generate a plausible looking record, write a rule against it, watch the rule fire, and ship it. The real device sends something slightly different. The rule never fires again and nobody notices, because a rule that does not fire looks exactly like a quiet network.
Here are the places I got this wrong, or nearly did.
CEF that is not CEF
SentinelOne was the best example. Its current syslog output opens with CEF:2|SentinelOne|Mgmt| and then abandons CEF entirely. What follows is not the standard header triplet of signature ID, name and severity. It is pipe separated key equals value pairs, with eventID, eventDesc and eventSeverity appearing as ordinary fields somewhere in the middle of the line.
If you write a CEF decoder and point it at that, it will parse the first three pipes and then produce nonsense. This is exactly why Wazuh ships no stock decoder for SentinelOne and why people have had to hand write them.
There were more quirks underneath. The product misspells one of its own field names as sourceMgmtPrecievedAddress. It emits Python list literals inside values, so you get a bracketed list sitting in the middle of a key value line. Backslashes in file paths are not escaped. And eventSeverity is 1 on every single record regardless of how serious the event is, because severity actually lives in the syslog priority.
LogGen reproduces all of that, misspelling included. A tidied up version would be easier to read and would not match anything real.
Positional CSV, where order is everything
Palo Alto PAN-OS writes comma separated values with no keys at all. A TRAFFIC log has 53 fields. A THREAT log has 60. Which field is the source IP depends entirely on counting commas, and several positions are literally called FUTURE_USE and exist only as placeholders.
Get one field out of position and every field after it is wrong. Your rule matching on a destination port is now reading a session ID. It will still compile, still run, and never match.
There is no way to guess this. You either have the field order from the vendor or you do not have a working source.
The same vendor, two incompatible formats
Oracle taught me this one. Oracle writes audit records in two different shapes depending on whether the database uses standard auditing or unified auditing:
Oracle Audit[pid]: LENGTH: "n" NAME:[len] "value"
Oracle Unified Audit[pid]: LENGTH: 'n' NAME:"value"Look at the quoting. Double quotes against single quotes, a bracketed length prefix in one and not the other. They are not interchangeable, and a decoder written for one silently fails on the other. Both are in LogGen as separate controls, because a real estate might be running either.
The other Oracle trap: ORA- error codes appear in the alert log, not the audit trail. Putting them in the wrong file would produce records that look right to a human reading them and match nothing in production.
When the documentation will not load
This happens more than you would think. Trying to pin down Trend Micro Vision One, every vendor page holding the CEF field table either redirect looped or served a page saying the documentation had moved. The field table was simply not reachable.
What worked was going to the SIEM side instead. IBM publishes captured sample events in its QRadar DSM guides. Google SecOps publishes default parsers. Splunk add-ons and Elastic integrations ship pipeline fixtures with real data in them. Microsoft Sentinel publishes its parsers on GitHub. These are second hand, but they are second hand captures of real traffic rather than somebody reconstruction from a spec.
For CrowdStrike this turned out to be decisive in an unexpected way. The Falcon SIEM Connector can emit JSON, CEF or LEEF, and it was not obvious which one a syslog listener would receive. The answer was in the connector own configuration file: raw JSON is written only to a local file on disk and is never pushed to a syslog server. To forward over syslog at all you must select syslog output, which produces CEF. So a SIEM with a syslog listener receives CEF, and that is what LogGen generates.
Samples can be edited too
One published CrowdStrike sample looked authoritative and turned out to be hand tidied. The giveaway was internal inconsistency: the detection ID in one field did not match the same ID embedded inside a URL in another field of the same record. Somebody had redacted parts of it by hand and not kept them in sync.
That matters, because that sample showed single backslashes in file paths while the connector configuration and a separate raw capture both showed doubled ones. If you trust the edited sample you get the escaping wrong. LogGen emits doubled and says in the source file that this is the one thing to check against a real forwarder.
Writing down what you do not know
The habit I ended up with is this. Every source file in LogGen has a comment at the top that cites the documentation it was built from and then splits into two lists: what was confirmed, and against what source, and what remains inferred.
For AWS CloudTrail, for instance, the record envelope is confirmed field by field against AWS documentation, including its own casing inconsistency where requestID has a capital ID and recipientAccountId does not. But the request parameter bodies for about twenty API calls that AWS publishes no CloudTrail example for are inferred from each service API reference. That is written down.
An admitted gap is useful. Somebody reading the file knows not to write a production rule against that field yet. A confident guess is a liability, because it looks exactly like knowledge.
The generator bug that keeps coming back
One last thing, and it is a coding problem rather than a research one. If a value appears twice in a record, generate it once.
This has bitten me repeatedly:
A Windows 4648 event where the IP in the message text did not match the IP in the event data.
A launchd persistence record naming one job label with a plist path pointing at a different one.
sc.execommand line on a SentinelOne tamper record that namednet.exe.A mitigation record classifying
mimikatz.exeas ransomware.
Each one compiles. Each one looks fine at a glance. Each one is a record no real device would ever produce, and any rule correlating those two fields would break on it.
The short version
Research the format before writing the code. Prefer a captured sample over a specification. Be suspicious of samples that contradict themselves. Write down what you could not verify. Generate a value once when it appears twice.
LogGen is open source under MIT at github.com/theshahrukh98khan/LogGen. If a format in there does not match what your own devices send, that is a bug worth telling me about.
Related posts
- I got tired of waiting for attacks to test my detection rules
You write a Wazuh rule, and then you wait for something to trigger it. Maybe that happens this week. Maybe the rule sits there with a typo in it for two months. LogGen is what I built to close that gap.
- Testing a Wazuh rule end to end without waiting for an attack
Prove the pipe before you trust a rule, remember that Wazuh does not listen for syslog by default, read the archives log before the alerts log, and always test the negative case. Most of the time lost onboarding a log source goes on not knowing which half of the chain is broken.
- T1047-Windows Management Instrumentation
MITRE ATT&CK Technique: T1047-Windows Management Instrumentation. Detections, visibility, use cases and real world attack insights.