AUR Diff Sentinel
A local-first CLI for reviewing AUR package updates and surfacing security-relevant changes.
Overview
What it is
AUR Diff Sentinel is a personal, local-first command-line tool designed to assist with the manual review of updates from the Arch User Repository (AUR).
It integrates into my update workflow by reviewing the diffs and PKGBUILD files of pending AUR package updates, while also supporting direct scans of local PKGBUILD files and unified diffs. Security-relevant patterns and changes, referred to as findings, are reported along with their location, the reason they deserve attention, and a review severity, allowing the most relevant changes to be manually inspected first.
A fundamental principle of the project I decided to adopt is that AUR Diff Sentinel is a triage tool, not a malware detector. A finding does not confirm a threat, nor does the absence of a finding guarantee safety. The tool surfaces evidence and review points that may warrant further investigation, while the final review and the decision of whether or not to update always remain with the user.
The goal is to make manual AUR review more focused, not to replace it.
Motivation
Why I built it
The idea behind the creation of AUR Diff Sentinel stems from a mix of habits and mistrust.
In 2026, there were several waves of malicious activity in the AUR, prominently involving the adoption of orphaned packages followed by malicious updates being pushed to them. As an Arch Linux user, this reinforced something I already knew but tended to overlook over time: installing software from the AUR, the Arch User Repository, requires placing a certain amount of trust in package maintainers, upstream projects, dependencies, and so on.
I was already in the habit of glancing at update diffs and PKGBUILD files, but these checks were often quick and superficial. It’s easy to do a poor job when conducting actual reviews of various updates, for every single update. Small changes can easily go unnoticed amidst version bumps, checksum updates, dependency changes, and the many other routine modifications involved in package maintenance. The more an update included these minor details, the more familiar it looked, and the more likely I was to gloss over something that actually deserved closer scrutiny.
Given my cybersecurity paranoia, I felt the need for an additional layer in my workflow to assist with the review process (without replacing me, of course). I didn’t want a quick manual inspection to be the only thing standing between an AUR update and my system, but I also didn’t want to delegate the decision of whether an update should be trusted to an automated tool. Having an assistant point me toward changes worth investigating felt like a useful middle ground.
This was the starting point for AUR Diff Sentinel: a tool built around my own update workflow, designed to make it harder to overlook security-relevant changes while keeping the final decision regarding trust firmly in my hands.
Design
How it works
AUR Diff Sentinel was designed around the idea that the two main workflow actions—reviewing and the actual update—were completely separate operations.
The tool now has two main entry points:
- A local
PKGBUILDor a unified diff; these can be scanned directly. - A pending AUR update. In this case, ADS queries helpers such as
yayorparuto discover available updates, then retrieves the corresponding AUR metadata for review.
pending AUR updates
↓
candidate metadata
↓
baseline ──→ diff
↓
analysis
↓
findings
↓
manual review
↓
external package update
↓
explicit baseline refresh
Baselines instead of trusting the latest state
Now, for the update review process, I needed a point of comparison. Therefore, AUR Diff Sentinel, which I will henceforth refer to as ADS for brevity, maintains a local baseline for each package undergoing review. You can think of this as a snapshot of the AUR metadata used as the previously accepted point of comparison.
When a new version becomes available, the relevant metadata is fetched separately and compared against the existing baseline. It is important to note that running the updates command does not replace the previous baseline with the one from the update. The candidate remains a candidate until the user has reviewed it and performed the package update.
If no baseline exists, for instance, during the initial use with a new package, the tool attempts to reconstruct the baseline from the AUR package’s Git history by locating the revision corresponding to the installed package. If that state cannot be retrieved, the current metadata can be scanned, but the tool explicitly warns that no update diff was available for review.
After an update is performed externally, the baseline refresh command checks whether the newly installed version matches the metadata that was just reviewed. If it does, the candidate effectively becomes the new baseline.
This ensures that the baseline represents a “last reviewed state” rather than merely a “last seen version.”
A layered analysis pipeline
I did not want the scanner to rely solely on a single detection strategy. While various patterns can be identified within a single line, many others only become significant when the surrounding context is considered.
The first scanning layer operates on individual lines of source code. Before rules are applied, lines are normalized into a common representation containing details such as the filename, line number, origin (file vs. diff), and, where possible, the execution context. This allows a rule to distinguish, for instance, between ordinary text containing curl and network activity occurring within prepare(), build(), check(), or package(). The same approach is used to identify contexts involving install scripts and pacman hooks.
Unified diffs receive special handling. Added lines are scanned for suspicious commands and patterns, while the “old” side of the diff is preserved for comparisons where examining a single line in isolation would be insufficient.
Next, a second layer analyzes changes in package metadata as distinct modifications rather than isolated strings. This layer can compare elements such as:
- source URLs and their domains
- changes from HTTPS to HTTP
- checksum arrays and algorithms
- newly introduced
SKIPvalues - mismatches between source and checksum counts
- dependency additions, removals, and group changes
- newly added install scripts, hooks, executable scripts, or binaries
This is particularly important for changes that only become significant when compared to their previous state. For example, a source domain cannot be considered “changed” without knowing where it pointed previously.
Looking at combinations, not only individual findings
Similar to what we said in the previous section, some signals become more “interesting” when they appear together.
For this reason, the analysis pipeline includes a correlation stage capable of combining various findings from the same review. For instance, a newly added dependency might require only moderate attention on its own; however, that same dependency appearing alongside an install script, weakened checksums, or network activity during a build could provide a different context for the review.
The scanner can thus track specific command sequences where necessary, rather than treating every line in isolation.
Naturally, I intentionally kept this correlation mechanism deterministic. The tool does not attempt to infer intent or determine whether a sequence is malicious; instead, it identifies combinations that warrant closer scrutiny during a manual review.
Findings as structured evidence
Now, how are all these conditions we talked about represented? Each of these detected conditions is represented as findings. A finding can bring with it the rule that produced it, its attention level, the file it comes from, the precise line, the matched content, the execution context and, for comparison, the old and new relevant values.
Keeping analysis separate from presentation makes the same evidence usable in different forms. The normal terminal report stays compact and groups findings by attention level, while --verbose can expose additional evidence and --explain can provide guidance about what was detected, why it matters, and what should be inspected.
The same underlying results can also be turned into a deterministic Markdown review packet, containing the relevant findings, package versions, files to inspect, and a rule-derived manual review checklist.
The reporting layer therefore changes how the evidence is presented, not what the scanner concluded.
Keeping the user in the loop
I won’t stress this enough, the whole workflow keeps the user at the center of the workflow, in fact this whole workflow stops before installing the package.
AUR Diff Sentinel uses AUR helpers for update discovery, but it does not wrap the actual update process!
After inspecting the findings, I still use my normal package-management workflow to decide whether and how to proceed.
The baseline is similarly never advanced simply because an analysis completed. If the installed version does not match the reviewed candidate, if analysis was incomplete, or if findings have not been explicitly accepted, the previous reviewed state is preserved.
That separation became one of the main design principles of the project:
discover
↓
analyze
↓
review
↓
decide
↓
update externally
↓
record the new reviewed state
AUR Diff Sentinel assists with the middle of that process, but it never collapses the entire chain into an automatic “scan and update” operation.
Implementation
Under the hood
AUR Diff Sentinel started as a very small line-oriented scanner, but by the time it reached 1.0 the implementation had grown into a set of small, deterministic analysis stages. I deliberately kept the runtime itself simple. ADS is written for Python 3.14+, uses a standard src package layout, and has no runtime Python dependencies.
Most of what the tool needs, like argument parsing, filesystem operations, diff generation, shell-like tokenization, subprocess management, regular expressions, URL parsing, and data modelling, is handled with the Python standard library. External programs are only used where they already represent the system state I want to inspect: paru or yay for AUR update discovery, pacman for installed package information, and Git for retrieving and walking AUR package history.
A common analysis model
One of the earliest implementation decisions was to avoid letting every check operate directly on arbitrary strings. The scanner revolves around a few small data structures:
raw input
↓
SourceLine
↓
Rule / structured analyzer
↓
Finding
A Rule describes a condition worth detecting, including its identifier, attention level, message, hint, and either a regular expression or a contextual matcher.
A SourceLine is the normalized representation of something the scanner can inspect. In addition to the actual text, it can carry the source filename, line number, whether it came from a normal file or a diff, its location in the diff and target file, whether it was added or removed, and any execution context that could be inferred.
A Finding is actual the concrete result. The finding preserves the rule that produced it together with the relevant evidence, like location, matched content, attention level, execution context and, when the finding comes from a comparison, the corresponding old and new values.
Keeping these objects independent from the terminal output became useful very quickly. Detection code can focus on producing evidence, while the reporting layer decides how that evidence should be presented.
Parsing files and unified diffs
Local files and unified diffs eventually converge on the same SourceLine representation, but getting there requires different parsing paths.
For a normal file, the scanner can simply walk the source line by line. For a unified diff, ADS first parses file headers and hunk ranges, keeping track of:
- old and new filenames
- hunk boundaries
- added, removed, and context lines
- line numbers inside the diff
- old-file and new-file line numbers
Only added lines are fed into the ordinary line-oriented rule scanner, because those are the lines introduced by the candidate update. Removed and context lines are still preserved where they are needed to reconstruct old and new metadata state.
On top of that, a lightweight context tracker follows the functions normally relevant to a PKGBUILD:
prepare()
build()
check()
package()
It also recognizes install-script functions and pacman-hook files. This lets the same piece of text carry different meaning depending on where it appears. A network command inside build(), for example, can be treated differently from the same word appearing in unrelated metadata.
The context tracker is intentionally small. It follows the function boundaries needed by ADS; it is not an attempt to parse or execute Bash.
Parsing just enough PKGBUILD syntax
Some checks cannot be implemented reliably by looking at independent lines. Sources, checksums and dependencies are commonly represented as Bash arrays, often split across multiple lines.
For these cases I implemented a small parser dedicated to the subset of PKGBUILD syntax that ADS actually needs. It understands supported source, checksum and dependency arrays, including architecture suffixes and += assignments, and collects their values while preserving filenames and source locations.
This allows the scanner to reason about structures such as:
source_x86_64=(...)
sha256sums_x86_64=(...)
depends=(...)
makedepends=(...)
rather than treating their contents as unrelated strings.
The parser uses shell-style tokenization where possible and falls back conservatively when the input is incomplete. .SRCINFO dependency changes are handled separately because its syntax is simpler and already normalized.
The same philosophy applies to shell commands. ADS contains lightweight helpers for splitting commands and pipelines while respecting basic quoting, stripping unquoted comments, ignoring common command prefixes such as environment assignments, and looking through bounded sh -c or bash -c payloads.
This is deliberately not a Bash interpreter. The implementation extracts enough structure to support the checks I actually use without pretending to understand every valid shell construct.
Building the detection pipeline
The actual scanner is composed of several kinds of analysis rather than one giant rule engine.
The first stage applies line-oriented rules. Some are regular-expression based, while others receive the full SourceLine and can inspect its execution context. This is where checks such as eval, shell indirection, downloaded content piped into a shell, network activity during build functions, live-system modifications, setuid/setgid permissions, and install-script or hook behavior are handled.
For complete PKGBUILD files, additional analyzers can inspect structures that span multiple lines. One example is checksum analysis, where a SKIP entry can be paired with the corresponding source and given different attention depending on whether that source is a VCS source.
Diffs receive another analysis path:
unified diff
↓
old/new metadata state
↓
structured comparisons
↓
findings
The diff parser reconstructs supported old and new array state and passes it to specialized comparison functions. Those functions handle source URL changes, domain changes, HTTPS-to-HTTP downgrades, removed or weakened checksum arrays, newly introduced SKIP values, source/checksum count mismatches, and dependency changes.
Dependency analysis covers both PKGBUILD and .SRCINFO. Added dependencies are classified conservatively, including special handling for build tooling, JavaScript tooling, and dependencies that appear likely to come from the AUR. The scanner also deduplicates equivalent dependency evidence when the same change appears in both metadata formats.
Finally, the correlation stage can generate composite findings from combinations that are more meaningful together than separately. It can also keep limited state across command sequences, for example, when package-manager activity occurs after moving into a temporary directory inside a live-system execution context.
The important implementation detail is that none of these stages produces a safety verdict. They all terminate in the same Finding model.
Implementing the update workflow
The AUR workflow adds orchestration around that scanner.
Update discovery is delegated to the tools already present on an Arch system. ADS runs paru -Qua or yay -Qua, parses the resulting package/version pairs, and then handles the review itself.
For each pending package, the candidate AUR repository is fetched through Git into the sentinel cache. ADS does not ask the helper to install, build, or execute anything.
The cache contains separate namespaces for the previously accepted comparison point and the latest reviewed candidate:
aur-diff-sentinel/
├── baselines/
│ └── <package>/
└── latest/
└── <package>/
By default this lives below $XDG_CACHE_HOME, falling back to ~/.cache/aur-diff-sentinel.
If a package does not yet have a baseline, ADS walks the AUR Git history and searches for a revision whose metadata version matches the version currently installed on the system. Version reconstruction accounts for epoch, pkgver, and pkgrel. Once found, that revision is copied into the baseline area and becomes the initial comparison point.
Candidate and baseline trees are then compared locally. The generated unified diff is passed through the same analysis pipeline used for direct diff scans, while newly added metadata files can also be inspected at the tree level.
Metadata reads are deliberately bounded and explicit. Symbolic links are not followed for analysis, oversized metadata files are rejected, invalid UTF-8 is reported, and unreadable files turn the affected package review into an incomplete analysis rather than silently disappearing from the result.
A failure affecting one package can normally leave the remaining packages reviewable. Failures that imply the cache itself could no longer be mutated safely are treated more seriously and abort the operation.
Cache mutation and recovery
Because the baseline represents review state, overwriting it carelessly would defeat the purpose of the workflow.
Cache-tree replacements therefore use a staged directory instead of writing directly over the current state. The new tree is prepared separately, the previous tree is moved aside, and the staged tree is published in its place. If the publication step fails, ADS attempts to restore the previous tree; if even that rollback fails, the backup is deliberately preserved and its path is reported.
The implementation also validates package names and checks that cache writes remain underneath the configured cache root.
This is not presented as a transactional database or as protection against every possible crash and concurrent writer. It is a small filesystem cache with explicit handling for the replacement failures the program can detect and recover from.
Interacting with the system
All external command execution is centralized behind a small runner abstraction. Commands are passed as argument vectors rather than interpolated shell strings, and each call has a fixed timeout.
The main interactions are intentionally narrow:
paru / yay -> discover pending AUR updates
pacman -> query installed package state
git -> fetch AUR metadata and inspect package history
pacman is also used when dependency analysis needs a conservative hint about whether a newly introduced dependency is unavailable from the official repositories and may therefore extend the AUR trust boundary.
This separation makes the same logic easier to test: command runners, metadata fetchers, installed-version lookups, and AUR-package checks can all be replaced with controlled implementations during tests.
Reporting the same evidence in different forms
The scanner never prints findings directly. The reporting layer receives the structured results and can render them in several ways. The default terminal view groups findings by attention level and keeps the output compact. --verbose exposes matched lines, contexts, hints and available old/new values, while --explain attaches short rule-specific guidance describing what was detected, why it matters, and what should be inspected manually.
For update reviews, the same result can be rendered as a deterministic Markdown review packet. The packet includes package versions, analysis status, findings, evidence, files identified by those findings, and a rule-derived checklist.
Package-controlled values are escaped for Markdown and are explicitly presented as untrusted evidence. The packet is therefore just another representation of the deterministic scanner output; generating it does not add a second decision engine.
The CLI also keeps operational state visible through its exit codes:
0 complete analysis, no findings
1 findings detected
2 operational error or incomplete analysis
That distinction is important for update reviews, because an incomplete analysis should not look equivalent to a clean one.
Testing the workflow, not only the rules
As the project grew, most of the difficult bugs were no longer simple cases of “this regex should match.” They appeared at the boundaries between package discovery, Git history, metadata parsing, cache state, partial updates, reporting, and failure recovery.
The test suite therefore uses synthetic package metadata, temporary directories, injected command runners and fetchers, and temporary local Git repositories. It does not need to query the live AUR or perform package updates.
Tests cover the individual parsing and detection layers, but also complete workflow properties such as:
- reconstructing baselines from historical metadata
- keeping candidate and baseline state separate
- refreshing only metadata that matches the installed version
- handling partial system updates
- continuing other package reviews after package-specific failures
- refusing to publish partially reconstructed baselines
- surfacing unreadable, oversized, invalid, or incomplete metadata
- preserving previous cache state when a replacement fails
- maintaining finding and exit-code behavior across normal reports and review packets
That testing strategy became particularly important during the hardening work before 1.0. The goal was not only to prove that suspicious examples produced findings, but also to make sure that failures in the surrounding review workflow stayed visible and did not accidentally turn into a false sense of completeness.
Security
Trust boundaries and limits
Security in AUR Diff Sentinel is deliberately framed around reducing uncertainty, not proving safety.
The tool exists because manual review can fail in very ordinary ways: a diff is long, a change looks routine, or something important is hidden between version bumps and checksum churn. ADS tries to make the parts that deserve attention harder to miss, but it does not attempt to answer the larger question of whether a package is ultimately trustworthy.
That distinction is fundamental to the entire project.
AUR Diff Sentinel helps me review the highest-risk changes first. It does not decide whether an AUR package is safe.
Security positioning
I intentionally describe ADS as a triage helper, not as an antivirus, malware detector, or formal verifier. A finding means that the scanner observed something matching one of the conditions I decided was worth reviewing. It may point to genuinely dangerous behavior, but it may also describe a completely legitimate packaging decision. Likewise, a clean report only means that no configured pattern was detected in the evidence ADS was able to analyze. It does not mean that the package is safe.
In other words:
finding detected
≠ malicious package
no findings
≠ safe package
analysis incomplete
≠ no findings
The HIGH, MEDIUM, and LOW levels follow the same principle. They express review attention, not a calculated probability that something is malicious.
A command that downloads and immediately executes remote content deserves more attention than a routine dependency change, but neither label is a verdict about intent.
What ADS tries to make visible
The scanner focuses on changes that can meaningfully alter where code comes from, what gets executed, or what a package is allowed to affect. At a high level, this includes changes involving:
- source URLs and source domains
- HTTPS-to-HTTP downgrades
- checksum removal, weakening, mismatches, or newly skipped verification
- install scripts and pacman hooks
- downloaded or decoded content being passed into a shell
- network access during build-related functions
- commands that may modify the live system
- setuid or setgid permissions
- shell indirection and simple forms of obfuscation
- dependency additions, removals, and changes in dependency role
- combinations of otherwise separate signals that deserve to be reviewed together
These checks are intentionally conservative. The goal is not to encode a definition of malware into a list of rules, but to identify events that change the review context. For example, changing a source domain may be perfectly legitimate because an upstream project moved infrastructure. That still changes part of the package’s supply chain and is therefore worth noticing. The same applies to dependencies. ADS can highlight that a new dependency appeared, including cases where that dependency may come from the AUR or introduces additional build tooling, but it does not recursively inspect that dependency and declare it trustworthy or untrustworthy.
The main trust boundary: package-controlled evidence
The most important security boundary in ADS is the boundary between package metadata and the rest of the software supply chain.
The AUR repository contains material such as the PKGBUILD, .SRCINFO, install scripts, hooks, patches, and other files committed alongside the package metadata. ADS treats that material as untrusted input and inspects it without relying on it to tell the truth about itself. Conceptually, the visibility of the tool looks roughly like this:
AUR package metadata
↓
baseline ↔ candidate
↓
AUR Diff Sentinel
↓
review evidence
↓
human review
But the complete path to the software that eventually runs on the machine is much larger:
AUR metadata
↓
upstream source / archive / binary
↓
dependencies and build tooling
↓
build process
↓
package installation
↓
runtime behavior
ADS has strong visibility over the first part of that chain and only indirect visibility over much of the rest. This matters because a perfectly ordinary-looking PKGBUILD can still reference an upstream project that has already been compromised. A downloaded archive can contain malicious code while matching the checksum intentionally published in the package metadata. A large source tree may hide behavior that is impossible to infer from the few lines changed in the AUR package. A checksum therefore answers a narrower question, whether the downloaded content matches the expected bytes, not whether those bytes are trustworthy. For the same reason, ADS does not attempt to treat a maintainer-controlled source URL and checksum change as automatically suspicious enough to prove compromise. If both are changed intentionally, the diff may still be internally consistent. The important security action is to notice the change and verify it against an independent source such as the upstream project’s official release information.
Package code is never part of the analysis process
One boundary I wanted to keep especially simple is execution: ADS does not execute AUR package code in order to analyze it.
It does not:
- execute a
PKGBUILD - build an AUR package
- run package install scripts or hooks
- install or upgrade packages
- execute downloaded package sources
The update workflow queries paru or yay only to discover pending updates, queries pacman for installed package state, and uses Git to retrieve and inspect AUR metadata. This means the scanner is analyzing text and metadata rather than observing a package by running it. That choice obviously limits what ADS can discover, but it avoids turning the review tool itself into another path through which untrusted package logic could execute.
The actual update remains outside the program. Only after I have reviewed the evidence and chosen to update through my normal package-management workflow can the reviewed candidate later become the new baseline.
Incomplete analysis must remain visible
A security tool can create a dangerous impression if a failure silently looks the same as a clean result. For that reason, ADS treats incomplete analysis as a separate outcome rather than converting it into “no findings.” Examples include metadata that cannot be read or decoded, files outside the supported analysis bounds, failures while retrieving package metadata, or errors while reconstructing an initial baseline. When this happens, the affected review explicitly carries analysis errors and the CLI reports an incomplete result. The stable workflow uses a distinct exit code for this state:
0 complete analysis, no findings
1 findings detected
2 operational error or incomplete analysis
This distinction also affects baseline management. A baseline cannot be refreshed from an incomplete analysis, even with --force. The purpose of --force is narrow: it lets the user explicitly accept known findings after making a manual decision. It is not a way to override missing evidence, failed analysis, or a mismatch between the installed package and the candidate that was reviewed.
I consider this an important property of the project:
Unknown evidence should remain unknown, not silently become clean evidence.
Conservative parsing, explicit limitations
ADS intentionally does not implement a complete Bash parser or shell interpreter. Instead, it extracts the limited amount of structure needed by its current checks: relevant PKGBUILD arrays, unified diff state, selected shell commands and pipelines, known build functions, install-script contexts, hooks, dependencies, and related metadata. This approach keeps the implementation understandable and deterministic, but it also creates unavoidable limits.
The scanner can miss behavior hidden behind shell constructs it does not understand. It can miss subtle logic spread across otherwise legitimate-looking code. Obfuscation more complex than the supported patterns may pass unnoticed. Generated, minified, binary, or patch payloads can contain behavior that ADS cannot meaningfully inspect. False positives are possible as well. Many operations considered security-relevant by the scanner are also perfectly legitimate in the right package. This is why a finding is always presented together with evidence to inspect, rather than converted into a binary classification.
What ADS cannot prove
There are several classes of problems that remain outside the security promise of the project.
ADS cannot guarantee protection against:
- an upstream project that is already compromised
- malicious source archives or binaries whose checksums match the values in the
PKGBUILD - malicious code hidden deep inside a large source tree
- subtle behavior expressed through valid but complex shell logic
- a maintainer intentionally changing both a source and its checksum
- payloads hidden inside generated, minified, binary, or patch files
- malicious behavior inside dependencies that ADS only identifies as dependency changes
- attacks that do not produce a change visible in the metadata being compared
- behaviors that fall outside the scanner’s supported syntax and detection rules
There is also a more general limitation: ADS sees evidence, not intent. The same technical action can be normal in one package and deeply suspicious in another. A new install script, a changed source host, or a package-manager command may have a legitimate explanation. Determining whether that explanation is credible requires context that cannot be reduced to a deterministic local rule. That final judgment remains part of the manual review.
The security promise
The security model of AUR Diff Sentinel is intentionally modest. It does not try to prove that an AUR package is safe. It does not assign reputation scores, recursively evaluate the trustworthiness of every dependency, or replace the need to understand what an update is changing.
What it tries to do is narrower and, for my workflow, more useful:
make risky-looking changes visible
↓
preserve the evidence
↓
surface uncertainty explicitly
↓
keep the final decision with the user
That is the boundary I want the project to maintain. A finding should make me investigate. A missing finding should never make me stop thinking (and being paranoid :D).
Outcome
Where it ended up
AUR Diff Sentinel started as a fairly small experiment: take a PKGBUILD or a diff, run a few security-oriented checks over it, and point out the lines that looked worth reviewing. It ended up becoming something much more complete than that initial scanner. By the time I reached the first stable release, the project had grown into an actual review workflow: update discovery, historical baseline reconstruction, metadata comparison, contextual and structured analysis, dependency review, deterministic reporting, explicit handling of incomplete analysis, and controlled baseline management all became part of the same tool.
What matters most to me, however, is that it stopped being only a project I was developing and became something I was actually using.
From prototype to daily workflow
A large part of the later development was driven by dogfooding. Instead of only feeding synthetic malicious examples into the scanner, I started running ADS as part of my own AUR update process. This exposed a different class of problems, like noisy output, confusing baseline states, partial updates, packages disappearing from the pending-update list after installation, recovery from failed metadata retrieval, and cases where the analysis technically worked but did not communicate clearly enough what the user should do next.
Before 1.0, I ran multiple complete validation cycles using the workflow as intended:
discover updates
↓
review with ADS
↓
manually inspect
↓
update externally
↓
refresh baselines
↓
verify state
Those cycles were particularly useful because they tested the assumptions between commands, not just each command independently. For example, the workflow had to continue making sense after only some of the reviewed packages had actually been updated, and baseline refresh had to work from cached reviewed metadata even when the helper no longer reported those packages as pending updates. That kind of behavior is difficult to notice when testing only the scanner in isolation.
Reaching 1.0
I treated 1.0 less as a feature milestone and more as a stability milestone. The goal was not to keep adding new detections indefinitely, but to reach a point where the existing workflow had clear and predictable behavior.
That meant hardening areas such as:
- baseline reconstruction and recovery
- package-specific failures without losing the rest of a batch review
- incomplete analysis remaining explicitly visible
- version identity, including
epoch,pkgver, andpkgrel - cache replacement and rollback behavior
- partial-update handling
- stable CLI commands and exit-code semantics
- deterministic review packets
- clean installation and documentation
The final release verification passed 162 tests, alongside CLI, help, package-installation, review-packet, and cache workflow checks. At that point, I was comfortable defining a stable 1.x contract around the parts of ADS that matter externally: the CLI, exit-code meanings, rule identities, and the safety properties surrounding baseline state. The exact wording of reports and the internal Python APIs remain free to evolve. The workflow guarantees are the part I care about keeping stable.
What changed in the way I approached the project
One of the main lessons from ADS was that the difficult part of a security tool is not necessarily detecting one more suspicious string. The project became much more useful when I started focusing on state, context, and failure behavior. A rule that detects curl | sh is easy to understand. Deciding what should happen when a baseline can only be partially reconstructed, when metadata cannot be decoded, when one package fails in a batch of ten, or when the user updates only half of the reviewed packages is much less visible, but just as important for whether the tool can actually be trusted as part of a workflow.
I also became much more careful about distinguishing three different states:
- nothing suspicious was detected
- something deserves attention
- I could not completely analyze this
Treating the third one as fundamentally different from the first became one of the most important principles in the project.
Another useful lesson was resisting unnecessary complexity. ADS uses lightweight parsing and deterministic analysis because that is enough for the problem I want it to solve. Where the scanner cannot understand something reliably, I would rather expose that limitation than quietly pretend to have deeper understanding than it actually does.
Where it stands now
Well, AUR Diff Sentinel reached its first stable release as a deterministic, local-first review assistant that I can integrate into my actual Arch Linux update workflow. It can discover pending AUR updates, reconstruct and maintain review baselines, compare package metadata, surface security-relevant changes, produce evidence-oriented reports, and keep incomplete analysis distinct from clean results, all without building, installing, or executing the packages it reviews.
I. Won’t. Stress. This. Enough. It still does not answer the question:
“Is this package safe?”
And I do not want it to.
The useful question has always been narrower:
“What changed here that I should look at more carefully?”
For me, reaching 1.0 meant that ADS could answer that question consistently enough to become part of the process rather than just an experiment beside it. The project is still active. There are ideas for extending the investigation workflow further, but any future layer has to preserve the same boundary that shaped the first stable release: deterministic evidence first, automation second, and the final trust decision left to the user.