Skip to content

docs(rfc): propose OpenShell testing strategy - #3460

Open
elezar wants to merge 6 commits into
mainfrom
codex/testing-target-state/el
Open

elezar wants to merge 6 commits into
mainfrom
codex/testing-target-state/el

Conversation

@elezar

@elezar elezar commented Sep 18, 2026 •

Copy link
Copy Markdown
Member

Summary

Propose RFC 0016 for OpenShell's testing strategy. Define separate ownership for behavioral contracts, test families, provisioning, installation, execution, and CI policy so contributors have one place to discuss the strategy and clear destinations for its living documentation.

Related Issue

Part of #3954, broadened at the maintainer's request to track the overall testing strategy and assign RFC 0016. Accepting the RFC does not close the implementation tracker.

Changes

  • Define general conformance, feature-specific, driver-specific, disruption, load/scale, and lower-level testing boundaries; retain SDK compatibility in companion RFC PR docs(rfc): propose SDK conformance testing #3238.
  • Define contracts, test cases, assertions, and suites; treat CLI and SDK access as client interfaces rather than separate test families.
  • Specify mandatory versus optional support and distinct result semantics; defer classification of capability-dependent portable tests until a concrete migration example.
  • Diagram target preparation separately from behavioral testing, reflecting tmachine environments, installers, and test suites.
  • Document existing conformance, Podman driver, and Keycloak provider-refresh suite locations and proposed additions.
  • Propose moving the shared conformance library under tests/suites/conformance after test(conformance): run driver suites with cargo #3866 removes the standalone conformance executable.
  • Assign clear responsibilities to TESTING.md, proposed tests/CONFORMANCE.md, suite READMEs, CI.md, and contributor skills.
  • Define incremental migration by behavioral intent and preserve unanswered versioning, admission, portability, reporting, and CI policy questions for review.

The final PR diff contains only the RFC. Current testing and CI reference documents will be updated through focused follow-ups as the strategy is agreed and implemented.

Testing

  • mise run pre-commit passes, including Markdown, Rust workspace/E2E/example lint, formatting, license, Helm, Python, and SDK checks.
  • Repository Markdown lint passes.
  • git diff --cached --check passes; the final diff against the PR merge base contains only the RFC.
  • RFC template sections and issue/PR references checked.
  • Unit and E2E tests are not applicable to this proposal-only change.

Checklist

  • Follows Conventional Commits.
  • Commit is signed off for DCO.
  • Living architecture and published documentation changes are deferred to implementation, as required by the RFC workflow.

@copy-pr-bot

copy-pr-bot Bot commented Sep 18, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@elezar elezar added the rfc label Oct 1, 2026
@elezar elezar changed the title docs(testing): define Nix and tmachine target state docs(rfc): propose OpenShell testing strategy Oct 1, 2026
elezar added 4 commits October 1, 2026 14:59
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
@elezar
elezar force-pushed the codex/testing-target-state/el branch from 5309c7e to 9c98a7f Compare October 1, 2026 13:08
elezar added 2 commits October 1, 2026 17:23
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
@elezar
elezar marked this pull request as ready for review October 1, 2026 16:05
@elezar
elezar requested review from a team, derekwaynecarr, mrunalp and sjenning as code owners October 1, 2026 16:05
| --- | --- |
| Unit and component integration | Internal logic and implementation mechanics, at the lowest effective layer. |
| General conformance | Public behavioral contracts across drivers and environments. |
| Feature-specific | Features requiring configured external integration or currently implemented on only one driver. |

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If something is only implemented on one driver, I'd assume that should be part of "driver specific"


| Family | Purpose |
| --- | --- |
| Unit and component integration | Internal logic and implementation mechanics, at the lowest effective layer. |

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this only unit and integration? is everything else an e2e test then?

| Family | Purpose |
| --- | --- |
| Unit and component integration | Internal logic and implementation mechanics, at the lowest effective layer. |
| General conformance | Public behavioral contracts across drivers and environments. |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is general conformance different than feature-specific where all drivers support it? I'm imagining that there's a set of features w/ venn diagrams including specific drivers, and in a world where all drivers are included, it just becomes a circle and is now deemed a "conformance test"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

effectively wondering if conformance is just a special-case version of feature-specific tests where all drivers are included, or if you see it as something different

Tests must not change gateway startup configuration. They may mutate public
API-managed state, using unique names and cleanup and avoiding conflicting
global-setting changes within a run. Document global effects; restoring prior
global settings is recommended, not mandatory.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

at a practical level I don't see any issues with global-setting changes if we're isolating tests, but I find this statement a little confusing. is it implying that multiple tests are sharing the same environment, so a good test-citizen should restore the env how they found it?

distinguish test corrections from changes to promised behavior, including
withdrawal of advertised support.

### 3. Make applicability and coverage explicit

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

making sure I understand this section:

Rather than encoding in our test-suite which drivers support features a/b/c, that should be encoded as an API contract that consumers can see. The the test suite just becomes one consumer of that API, and decides which tests are worth running against this given openshell deploy?


| Family | Purpose |
| --- | --- |
| Unit and component integration | Internal logic and implementation mechanics, at the lowest effective layer. |

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

its a little confusing to put unit and integration in the same sentence. these are usually distinct types of tests and integration has become very overloaded. we call current conformance tests, "conformance integration", and current feature specific tests "feature specific integration"

Image

and a suite groups related cases. For example, deleting a sandbox must remove
it from the sandbox list; an assertion checks that its identifier is absent.

| Family | Purpose |

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would suggest the following taxonomy

Test family Description
Lint Check formatting, style, and static rules.
Unit Verify one component in isolation. Place tests inline in its Rust module or in an adjacent test.rs file under src/.
Integration Verify interactions between components. Place Rust integration tests in the crate’s top-level tests/ directory.
End-to-end Verify a configured OpenShell target through a client interface. Place suites in the repository-level e2e/ directory; conformance is a class of end-to-end test for portable public contracts.
Benchmark Measure performance or scale. Place crate benchmarks in benchmarks/ and full-system benchmarks in a root benchmarks/.

Types of e2e tests

End-to-end type What it verifies
Conformance Portable public contracts across applicable drivers, including advertised optional capabilities. Includes recovery and continuity after an induced failure. It should be possible to run these tests out of tree.
Feature Behavior requiring a named external service or special gateway configuration.
Driver Behavior specific to a driver, its host integration, or its configuration.
Installation Installing, upgrading, and uninstalling candidate artifacts.

Comment on lines +45 to +48
A behavioral contract specifies an operation's expected observable result under
stated conditions. A test case exercises it, an assertion checks an observation,
and a suite groups related cases. For example, deleting a sandbox must remove
it from the sandbox list; an assertion checks that its identifier is absent.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems unrelated to the section.

suites. Installation, upgrade, and uninstall behavior need separate assertions;
successful conformance alone does not validate packaging.

## Implementation plan

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be good if we could start to define what conformance tests we're going to build. I would propose the following suites. Here's a list to get us started

  • Sandbox lifecycle: Create, inspect, execute, stop, restart, and delete.
  • Policy: Validate, apply, update, and report effective policy.
  • Sandbox enforcement: Enforce filesystem, process, and network rules.
  • Providers: Manage providers and verify credential delivery, isolation, rotation, and protection from exposure.
  • Identity and authorization: Authenticate and enforce access boundaries.
  • Middleware: Verify selection, ordering, transformation, and failure handling.
  • Interceptors: Verify request transformation, rejection, and preservation of authorization.

├── config.nix # tmachine definitions
├── artifacts.nix # Artifact construction
├── ansible/ # Provisioning and execution
├── CONFORMANCE.md # Proposed: agreed policy

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why is conformance at the top instead of inside conformance/?

specialized infrastructure. Performance thresholds belong in load/scale unless
the deadline is itself a public contract.

### 2. Define conformance through public behavior

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this whole section is pretty dense. i'm not sure i understand it. conformance tests should

  • use public apis
  • be parametrized by gateway endpoint
  • have the ability to run out of tree so third parties can test for conformance
  • run as part of nightly qualification
  • run on branch checks when manually triggered.
  • it would be great if we can granularly trigger conformance checks by suite on ci

configuration, or equivalent target identity. Record mock targets as mocks,
not evidence for production drivers.

### 4. Separate target preparation from test execution

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we please detail what tests run where. for example, unit tests always run on ci, e2e tests run when manually triggered or on release qualification, etc. we also need to cleanup the current labelling approach.

i would also like to see us be more efficient in what tests run on branch checks. @SDAChess mentioned work to dynamically figure out what tests to run per pr. i think we should consider this. i've seen interesting approaches that use llms or jev to figure out what tests should be run based on changes. might be interesting to experiment with.

we don't have to build this all at once, but since this rfc is broadly titled "propose OpenShell testing strategy" and talks about execution stragies, i think we should detail where we're headed with things. current branch checks on prs are too slow.

Comment on lines +119 to +120
| Unsupported | Optional support was not advertised. |
| Skipped | The test was deliberately excluded. |

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i'm assuming this is just for conformance tests. how do we specify a test to be unsupported or skipped?

Comment on lines +104 to +108
The gateway must report effective capabilities for its running configuration
through the public API and machine-readable CLI output. Discovery failure aborts
conformance. Start with flat, namespaced booleans; defer hierarchy, parameters,
and profiles. Capabilities describe product behavior, not test selectors,
credentials, or external-service prerequisites.

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure how scalable this is. For example, what if down the road we support loading more than one compute driver. Some compute drivers might enable different capabilities. This is also going to cause options on the compute driver to explode with every product capability a driver might or might not support. We've already started to see this and it creates quite a change amplification that I'd like to avoid.

As part of RFC-0012 we proposed using validation to assert capabilities. For example, if you create a sandbox with a file system policy, and no filesystem policy exists, that sandbox should fail to create will a validation error. I think this is a more scalable approach.

Comment on lines +97 to +100
Use normal PR review, without a soak period or separate promotion PR. Resolve
known flakiness rather than hiding it with retries. Changes or removals must
distinguish test corrections from changes to promised behavior, including
withdrawal of advertised support.

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should handle some level of flaky-ness and include infra for retries. If a test is identified as flaky there should be a report that we monitor and fix out of band. If we hard fail on flakes, we're going to be fighting builds and wind up just manually retrying tests anyways (we already see this today).

Down the road, if there's a report we can have some agent iterate on the report to reduce flakes.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants