Skip to content

Inside OBJECTS.DATA: WMI Repository Forensics

How the WMI CIM repository stores objects: OBJECTS.DATA pages, INDEX.BTR keys, the three MAPPING files, class definitions, instances and name hashes.

Published on 4 min read

TL;DR. The WMI repository is a small database in C:\Windows\System32\wbem\Repository. OBJECTS.DATA holds records in 8 KiB pages, INDEX.BTR is a B-tree of text keys built from SHA-256 name hashes, and the three MAPPING files turn logical page numbers into physical ones. Class definitions describe the layout; instances hold the values. Freed pages keep their old bytes, which is what makes recovery possible.

The layout below comes from FireEye's 2015 research and the python-cim parser by Willi Ballenthin; the encoding inside each record follows Microsoft's public [MS-WMIO] specification. Microsoft does not document the repository container itself, so treat field names as the research community's.

The files

FileRole
OBJECTS.DATARecords: class definitions and instances, 8 KiB pages
INDEX.BTRB-tree index, 8 KiB pages, keys point into OBJECTS.DATA
MAPPING1.MAP, MAPPING2.MAP, MAPPING3.MAPLogical → physical page maps for both files, three generations

On Windows XP and Server 2003 the same files live in Repository\FS\ and names are hashed with MD5 instead of SHA-256.

Pages and the table of contents

A page of OBJECTS.DATA that starts records opens with a table of contents: 16-byte entries (record id, offset in the page, size, checksum) closed by an all-zero entry. A record larger than the rest of its page continues on the following logical pages, which have no table of contents.

Logical pages are what the index talks about. Where a logical page physically sits is the job of the current MAPPING file: each entry gives the physical page number of one logical page, and a special value marks unused ones. Physical pages that no entry references are free — and still hold whatever was written there last.

The index: keys built from hashes

INDEX.BTR stores text keys. Each part is a prefix plus the upper-case hex SHA-256 of the upper-cased name in UTF-16:

NS_<h("ROOT\SUBSCRIPTION")>/CD_<h("COMMANDLINEEVENTCONSUMER")>.268.2079887499.561
NS_<h("ROOT\SUBSCRIPTION")>/CI_<h("__EVENTFILTER")>/IL_<h(key)>.271.2079886610.299

NS_ is the namespace, CD_ a class definition, CI_ + IL_ an instance of a class. The suffix is logical page . record id . length. With the current map, that is enough to fetch any object — which is what "structured mode" means in WMI Parser.

Because names are hashed, the index never contains a readable namespace or class name. Namespace names come from __NAMESPACE instances (each namespace lists its children), and class names from the class definitions.

Class definitions and instances

A class definition record holds the superclass name, a FILETIME, then the class part described in [MS-WMIO]: qualifiers, a property table (name, CIM type, declaration order, offset in the value table) and the class defaults. __EventFilter and __FilterToConsumerBinding are defined in the __SystemClass namespace; the standard consumers in root\subscription.

An instance record starts with:

  1. the class name hash, as 64 UTF-16 characters (32 on XP);
  2. two FILETIMEs, undocumented (python-cim calls them timestamp1 and timestamp2);
  3. the instance part: a 2-bit-per-property table saying whether each value is set, NULL or the class default; a fixed-size value table; qualifiers; then a heap where strings and arrays live.

Strings in the heap are stored as a flag byte (0 for 8-bit text, 1 for UTF-16) followed by NUL-terminated text. That is why a plain text search in OBJECTS.DATA finds CommandLineEventConsumer.Name="…" and the command lines themselves.

Decoding an instance needs its class layout: the property list comes from the class definition and all its ancestors (CommandLineEventConsumer → __EventConsumer → __IndicationRelated → __SystemClass).

What this means for an investigation

  • Live objects are the ones reachable through the index and the current map.
  • Deleted objects stay in unmapped pages or in the unused tail of a page until overwritten. Scanning for class-name hashes finds them; the MAPPING file tells you they are no longer referenced. See recovering deleted WMI persistence.
  • All files must come from the same instant. An INDEX.BTR newer than the MAPPING file points at the wrong pages. Collect from a shadow copy: how to collect the WMI repository.
  • Timestamps are weak. Class definitions carry a FILETIME and instances two undocumented ones; none is a documented creation date.

WMI Parser implements both paths — the index walk and a full carve — and labels every object with how it was found.

FAQ

What is in OBJECTS.DATA?

Every class definition and class instance of every WMI namespace, stored as records in 8 KiB pages. Event filters, consumers and bindings are ordinary instances in it.

Why are there three MAPPING files?

Windows keeps three generations of the logical-to-physical page map so it can recover from a failed write. The one with the highest version is current; the others describe earlier states.

Can OBJECTS.DATA be read without INDEX.BTR?

Yes. Instance records start with the hash of their class name, so they can be found by scanning the file. The index adds namespaces and tells live records from deleted ones.

Related articles

Read CCM_RecentlyUsedApps from the WMI repository: which programs each user ran, when last and how often, including older copies carved from OBJECTS.DATA.
Why deleted WMI filters, consumers and bindings survive in OBJECTS.DATA, how index walking and carving differ, and how python-cim and PyWMIPersistenceFinder compare.
How WMI event subscription persistence works: __EventFilter, event consumers and __FilterToConsumerBinding, where they are stored and how to find them.