# Building AT1-Defender #2: Inside SCOPE

In the first article, we introduced **AT1-Defender**: why we started building it, what we want it to become, and the direction of the project.

This time, we are going deeper.

Before AT1-Defender can classify suspicious behavior, it needs reliable information about what is actually happening inside Windows. Processes start and terminate. Modules load. Files change. Registry objects are modified. Network connections appear and disappear.

Individually, most of those events mean very little.

Together, they can tell an entirely different story.

That requirement eventually became:

**SCOPE – System Correlation & Observation Processing Engine.**

SCOPE is our developing Windows telemetry and correlation engine, built to give AT1-Defender its own understanding of system activity.

With **SCOPE 0.2.0.0**, that project has moved far enough that we can finally talk about what it actually does, what broke while building it, and how close it is getting to becoming something much more serious.

## Before SCOPE

During the earliest development of AT1-Defender, we needed system telemetry long before we had the architecture required to collect it ourselves. Sysmon became an important development tool, allowing us to observe processes, network activity and other Windows events while we worked on the protection systems above them.

But we realized relatively early that **we did not want the final AT1-Defender product permanently dependent on Sysmon for its behavioral visibility.**

Part of that was architectural. We wanted greater control over process identity, correlation, event handling, recovery and how incomplete information was represented.

There was also a security consideration. Established monitoring technologies are extensively documented and understood. That is enormously valuable to defenders, but those same technologies are inevitably studied by malware developers. Modern threats can inspect their environment, identify monitoring components and sometimes change their behavior when they believe they are being observed.

That does not make Sysmon insecure, and writing our own code certainly does not make us magically invisible. The issue was control.

If telemetry would eventually influence AT1-Defender's security decisions, we wanted to understand and control much more of the path between **Windows doing something** and **AT1-Defender understanding what happened**.

Simply recreating Sysmon would therefore accomplish very little.

The question changed from:

> **How do we replace Sysmon?**

to:

> **If AT1-Defender owned its system visibility, how should that system actually work?**

That experiment gradually developed its own event model, process identities, correlation logic, processing pipeline and health monitoring.

It became **SCOPE**.

## Seeing Windows Was the Easy Part

The first prototypes concentrated on the fundamentals: process activity, filesystem events, Registry operations, module loading and network connections.

Getting events to appear was encouraging.

Understanding them was harder.

Imagine that Windows reports a process starting, a file being created, a Registry value changing and a network connection appearing shortly afterwards.

A logger can simply record four events.

SCOPE needs to determine whether those events actually belong together.

That distinction quickly became fundamental to the architecture.

**Observation is not interpretation.**

A `.dll` being accessed does not prove it was loaded as a module. A Registry operation does not necessarily mean that a modification succeeded. Two events occurring milliseconds apart do not prove that one caused the other.

If SCOPE cannot establish something reliably, we would rather preserve that uncertainty than manufacture a more impressive-looking answer.

And nowhere did that become more obvious than process identification.

## The PID Problem

A Windows Process ID looks like an identity.

It isn't.

Windows reuses PIDs. A process can terminate and another completely unrelated process can later receive the same number.

That creates an interesting problem for telemetry.

Suppose process `6420` performs an operation and terminates. Windows later assigns `6420` to another process, while delayed telemetry from the original process is still moving through the system.

A naïve correlation engine could attach the event to the wrong process.

Nothing necessarily crashes.

Everything could appear perfectly healthy.

**SCOPE would simply be telling the wrong story.**

That was unacceptable.

SCOPE therefore moved toward lifecycle-aware process identities. Events are correlated against the process instance that existed when the activity occurred rather than blindly trusting a PID.

The same philosophy applies when information conflicts. If SCOPE cannot safely determine a parent relationship or attribution, the result can remain unresolved.

**Unknown is better than confidently wrong.**

## So We Tried to Make It Blind

Once basic collection worked, we stopped concentrating on adding features.

Instead, we started attacking what we already had.

Processes were rapidly created and terminated. Events were reordered. PID reuse was simulated aggressively. Filesystem activity increased. Collection was interrupted. Queues were stressed. We deliberately created situations where a process disappeared before all associated telemetry had finished travelling through the pipeline.

And we found problems.

Some event ordering could influence attribution. Process relationships could occasionally become more certain than the evidence justified. Historical information could become dangerous when identifiers were reused.

Those failures changed both SCOPE and the way we tested it.

A successful test could no longer mean simply:

*"We received the event."*

It needed to mean:

*"We received the correct event, associated it with the correct process, and knew when that association could no longer be trusted."*

Every reproducible failure began becoming a regression test.

Because fixing a correlation bug once is useful.

Making sure it stays fixed is considerably more useful.

## When the Pipeline Became the Problem

There was another problem waiting for us: volume.

Windows can produce telemetry considerably faster than an endpoint security product should synchronously analyze it. During early endurance testing, we found a bottleneck in the way diagnostic telemetry was being retained.

Instead of simply increasing buffers, we changed the architecture.

SCOPE moved toward bounded processing and continuous batch delivery. Incoming telemetry can continue through the pipeline without requiring an entire session to accumulate in memory, while heavier processing stays away from the time-sensitive collection path.

That produced another rule:

**If SCOPE cannot keep up, SCOPE must know that it cannot keep up.**

Silently losing events while continuing to report complete visibility would be far worse than openly reporting degraded telemetry.

The same applies when collection itself fails.

SCOPE now monitors the health of its own telemetry pipeline and can recover from certain interruptions without requiring AT1-Defender itself to restart. More importantly, a successful recovery does not erase the fact that visibility was temporarily interrupted.

Recovery restores visibility.

It does not rewrite history.

## Proving It

By **SCOPE 0.2.0.0**, watching events appear in a development console was no longer enough.

We needed numbers.

One of our attribution tests repeatedly reordered **1,000 process lifecycles using the same simulated PID**. Across those scenarios, SCOPE evaluated **20,000 observations with known ground truth**.

All expected observations were attributed correctly. Events without sufficient evidence were deliberately left unresolved.

We then performed **20 repeated live verification runs**, checking process relationships, module activity, file creation, Registry activity and TCP destinations.

Those tests passed.

A separate overload test pushed **100,000 events** through parts of the pipeline, while resource measurements gave us a baseline for memory, CPU usage, queue pressure and processing latency.

Then came one of the more interesting live tests.

An earlier run had exposed a bottleneck in telemetry delivery. After redesigning that part of the pipeline, SCOPE processed **11,762 ETW events during an approximately two-minute run without reporting event loss**.

The wording is intentional.

We are not claiming SCOPE cannot lose telemetry.

We are saying that under this particular measured workload, the revised pipeline processed those events without detecting loss.

That distinction matters.

The objective is not to produce impressive numbers. It is to establish baselines against which future versions can be judged.

For the first time, we could begin measuring not only what SCOPE observed, but **how well SCOPE itself was observing it**.

## Not Production-Ready. Yet.

SCOPE 0.2.0.0 has moved considerably beyond the original experiment, but we are not calling it production-ready.

Short tests cannot expose every slow memory leak or long-running synchronization problem. Controlled environments cannot reproduce every Windows installation. Additional telemetry sources will introduce new correlation problems, new performance costs and probably several bugs currently waiting patiently for us to discover them.

And there is something we are deliberately **not** doing yet.

SCOPE does not currently make autonomous destructive decisions from this telemetry.

We want the observation and correlation layers to earn our trust before AT1-Defender begins relying on them for increasingly consequential actions.

The development order is therefore intentional:

**Observation → Correlation → Detection → Enforcement.**

Skipping directly to the last step would certainly produce a more exciting demonstration.

It would also be an excellent way to delete something important for entirely the wrong reason.

## Where SCOPE Goes Next

There are already several directions we are investigating.

Additional network context, including DNS activity, could allow SCOPE to understand more of what happens before a connection is established. Service activity, scheduled tasks and additional persistence-related telemetry are also interesting candidates.

But we are deliberately avoiding a race to collect everything Windows can expose.

Every new sensor introduces additional processing, additional relationships and additional ways to reach the wrong conclusion.

The question is therefore not:

**Can we collect this?**

It is:

**Can AT1-Defender reliably use it?**

That is a considerably higher standard.

SCOPE began as an experiment to give AT1-Defender greater control over its Windows telemetry. It now has lifecycle-aware process correlation, bounded processing, recovery mechanisms, telemetry health monitoring and repeatable ground-truth testing. The current technical status remains a development candidate moving toward beta rather than a production-ready subsystem. Inklistrad markdown

The first versions proved that we could see Windows.

The next versions taught us how easy it was to misunderstand what we saw.

Now we are building the part that matters:

**learning when that understanding can actually be trusted.**

And that is where SCOPE starts becoming more than an experiment.
