Skip to content

Engineering & Evidence Standards

Technical work is most useful when visitors can tell what exists, what was measured, what remains limited, and who owns the final judgment. These standards explain the labels used throughout this site.

Maturity labels

Production, pilot, and R&D describe different kinds of evidence. None should be promoted into another category by implication.

Production

A system is deployed into a real operating environment with active ownership. The label does not imply unmeasured scale, uptime, adoption, or business impact.

Pilot

A bounded workflow is being tried with real users or operating context, but broader rollout, long-term reliability, and adoption are not yet established.

R&D / Prototype

A technical, academic, or active-build artifact tests an approach. It can demonstrate engineering decisions without being presented as paid deployment or mature operations.

What counts as evidence

  1. 01Production systems with real operational ownership
  2. 02Measured business, user, quality, or reliability outcomes
  3. 03Reviewable architecture, code, deployment, and platform artifacts
  4. 04Enterprise engineering history with bounded responsibilities
  5. 05Published research, reproducibility, and accepted work with status shown
  6. 06Credentials, memberships, and technology lists as supporting context

Limitations and claim boundaries

Prototypes keep their prototype label until deployment and operating evidence justify a change. Missing screenshots, usage data, monitoring, or outcome measurements remain missing rather than being replaced with adjectives.

Limitations, failure modes, data coverage, refusal behavior, and ownership boundaries are disclosed when they materially change how a result should be understood. Unsupported scale, accessibility, reliability, and business-impact claims are avoided.

Architecture is described as a capability and responsibility. Historical job titles stay tied to the employers and records that established them.

Research status and reproducibility

Publication status is evidence too, so accepted work and formally published work are kept separate.

Published research

A work is presented as published only when a formal publication record supports that status. DOI, byline, affiliation, venue, and date are checked where available.

Accepted research

Accepted work is labeled accepted until proceedings or another formal publication record is available. Acceptance is meaningful, but it is not described as publication.

Evidence quality

Separate measured results, accepted work, prototypes, and production proof.

Retrieval evaluation

Test whether the system finds and grounds the material a workflow depends on.

Limits and uncertainty

Document coverage gaps, failure modes, uncertainty, and refusal boundaries.

Reproducibility

Keep data preparation, evaluation conditions, and technical decisions reviewable.

Scientific datasets

Bring careful processing habits to noisy, incomplete, and domain-specific data.

AI-assisted engineering and human review

Assistance can accelerate the work, but it does not transfer engineering accountability.

AI tools may support research, planning, code generation, testing, documentation, or review. A human remains responsible for requirements, source verification, design decisions, code review, security and privacy judgment, test interpretation, deployment, and the claims made about the finished work. Generated output is treated as material to inspect, not evidence by itself.

Quality evidence

Evidence is matched to the property being claimed and kept within the boundary of the recorded check.

Security

Access boundaries, data handling, dependency checks, threat considerations, and remediation are cited only to the level demonstrated by the work.

Accessibility

Semantic structure, keyboard use, labels, contrast, and automated or manual checks are treated as engineering evidence. Scores are not claimed without a recorded run.

Testing

Unit, integration, end-to-end, and behavioral evaluation are selected for the actual failure risks. A passing test is not generalized beyond what it checks.

Reliability

Failure handling, observability, recovery, deployment discipline, and operational ownership matter. Uptime or performance claims require measurements.

Ask about a specific claim or engagement