Back to Blog

Rules of Engagement and Production Safety for Institutional Blockchain Penetration Testing

Code Auditing
3 de setembro de 2026
10 min read
Key Insights
  • A useful penetration-testing engagement is prepared before it runs: a clear objective, agreed scope, accountable owners, and authorized access, with a plan to keep normal operations throughout—preparation that reflects how the institution moves and controls value.

  • A Rules of Engagement (RoE) document makes the engagement executable by setting authority, activity limits, communications, escalation, and evidence handling.

  • Production testing requires service safeguards, measurable stop criteria, monitoring, change coordination, and pause authority; the engagement should also define remediation and retest so findings become validated improvements.

The first two articles in this series set out why crypto institutions need blockchain penetration testing (Part 1) and what the discipline is (Part 2). This article turns to how an institutional engagement is prepared and run safely: the rules of engagement, the production-safety safeguards, and the remediation and retest that turn findings into fixes—the engagement lifecycle shown below.

Figure 1. Institutional penetration testing engagement lifecycle.
Figure 1. Institutional penetration testing engagement lifecycle.

Web3 raises the stakes of production testing in a specific way: on-chain actions are typically irreversible, test activity is often publicly visible on-chain, and the systems in scope can move real funds. The safeguards below therefore weight environment choice, value limits, key handling, and reconciliation more heavily than a conventional engagement.

Blockchain Penetration Testing

Find the path in — across contracts, nodes, APIs, and cloud

RoE as the operating authority

According to the U.S. National Institute of Standards and Technology (NIST), RoE sets the guidelines and constraints for security testing [1]. For an institutional engagement, it turns a decision to test into an authorized mandate: a defined objective, a defined scope, and named authority to act.

That mandate matters because a test team cannot safely infer authority from a contract or asset list. Without a shared decision about the risk being assessed, the systems involved, and the people responsible for them, the team may test the wrong path, exclude a critical dependency, or lack the authority to validate a material finding.

The starting point is the business decision that the test is intended to support. An objective may concern public exposure before launch. It may focus on customer and application programming interface (API) authorization, or on the reachability of privileged cloud and operational systems. It may also test a signing or withdrawal workflow, or assess a third-party integration. That objective identifies the systems that matter, the risk to reduce, and the result management needs to make a decision.

The objective should translate into a scope that follows the operating model. In a crypto institution, this commonly includes customer-facing web, mobile, and application programming interface (API) services; cloud accounts and identity and access management (IAM); continuous integration and continuous delivery (CI/CD) systems and secrets; operational consoles; wallets and approval workflows; signing systems; ledger and withdrawal services; and the vendors that connect them. Together, these components govern how customers and internal teams initiate, approve, sign, release, and reconcile fund-related actions. The institution and test team should also agree the relevant perspective: an external attacker, a normal user, a partner, a low-privilege employee, or an assumed-compromised identity.

Each system, account, interface, and activity in scope should have a named owner and a clear authorization path. The same applies to third-party dependencies, including custody providers, wallet interfaces, remote procedure call (RPC) providers, software as a service (SaaS) platforms, identity providers, code hosts, and managed services. Testing a provider-owned environment requires that provider's written permission; the institution's authorization alone may not authorize activity against the provider's systems.

Together, the objective, scope, and authorization path establish what the team may assess. The next step records how the team may perform that assessment.

Operating bounds

Operating bounds are the written constraints that govern execution after the mandate is agreed. They distinguish what is in scope from what activity is allowed: a production withdrawal service may be in scope for authorization-path validation while real customer withdrawals, private-key extraction, persistence changes, and load-generating attacks remain prohibited.

This distinction prevents uncertainty from becoming operational risk. During a production engagement, the institution and test team need to know in advance which techniques are allowed, when activity must stop, who can make that decision, and how evidence may be handled. Otherwise, even an authorized test can create avoidable service or customer impact.

The RoE should record those decisions before testing begins. Moreover, the document should be precise enough for both parties to act without reopening fundamental questions during the engagement.

The table below translates the test mandate established above into a practical RoE record. It groups the decisions that must be settled before testing: what is being assessed, who is authorized, which activities and limits apply, how the parties coordinate, and how evidence is handled. It is not a generic checklist to copy unchanged; the recorded values should reflect the institution's operating model, test perspective, and production risk.

RoE topic Decision to record before testing
Objective and scope The business decision, systems and interfaces in scope, owners, and the attacker perspectives to be tested.
Authorization Written authority, test identities, approved access paths, and provider approvals for third-party environments.
Test method and completion Permitted techniques and the point at which the team must stop, such as demonstrated access, privilege escalation, or a controlled workflow simulation.
Operating limits Test windows, request-rate and concurrency limits, account-operation limits, data-access boundaries, change-freeze periods, and, where an approved scenario uses a transaction, the permitted network, test addresses, transaction types, maximum test value, and gas budget.
Prohibited activity Examples include denial-of-service testing, social engineering, real customer-asset movements, private-key export, persistence, or unapproved production changes.
Sensitive workflows Designated wallets and accounts, allowlisted destinations, maximum test value, approval participants, expected policy behavior, reconciliation steps, and the furthest point in the transaction-control chain that validation may reach.
Communications and pause Routine notice model, protected safety channel, escalation contacts, material-finding threshold, named authority to pause or resume testing, and the expected observability and alert handling for an approved public transaction broadcast.
Evidence handling Minimum necessary evidence, data minimization and redaction, encryption, approved recipients, retention period, destruction confirmation, and transaction-hash and network metadata linked to the internal approval and ledger evidence for an approved test transaction.

With the execution boundaries agreed, the institution can prepare the deployed systems and operating conditions that the test will encounter.

Operating environment and live-service protection

The operating environment is the deployed system and the business activity around it: trust boundaries, fund flows, service dependencies, operational workflows, and the people responsible for each component. It is the context in which a test finding acquires its actual meaning.

That context matters because the same technical weakness can have very different consequences. An API issue, cloud permission, or approval-workflow weakness may affect customer data, internal operations, signing authority, balance changes, or withdrawals. For a live exchange, payment company, custodian, or wallet provider, testing must also coexist with trading, payment, deposit, withdrawal, settlement, and support operations.

A current view should cover the asset and service inventory, architecture and integration points, cloud and identity model, operational workflows, and third-party dependencies. Planning should also identify relevant business windows, planned releases, change freezes, high-volume periods, hot-wallet activity, and other operational events. Service-health signals observed during testing should include transaction and API volumes, response times, error rates, queue depth, signing-service health, and wallet-service availability. The operating environment includes the production services, identities, operational workflows, and external dependencies used to initiate and control transactions. The RoE records the approved test environment, test identities, permitted steps within a signing or withdrawal workflow, and the safeguards that apply to any controlled validation scenario. These parameters keep the assessment focused on the institution's controls while protecting normal customer and operational activity.

The European Union's Digital Operational Resilience Act (DORA) provides a useful high-regulation reference point. For financial entities selected for threat-led penetration testing, Article 26 requires the test to cover critical or important functions and to be performed on the production systems supporting them, including relevant outsourced information and communication technology (ICT) services [2]. This is not a general authorization for production testing; each institution still needs its own authority, safeguards, and applicable legal and contractual approvals.

This production baseline lets the institution apply more specific safeguards to the workflows that can directly affect funds or customer access.

Boundaries for sensitive workflows

Sensitive workflows are the systems and actions whose normal operation can directly affect funds, customer access, or the integrity of records. They include signing and withdrawal workflows, privileged production access, customer-data handling, and ledger operations.

These workflows need stricter boundaries because a realistic test can otherwise cross from validating a control into changing a customer or financial outcome. The risk is not only technical disruption: it can include unauthorized fund movement, incorrect balances, exposure of sensitive data, or confusion between test activity and an actual incident. Where a control failure can affect funds, reversal may be difficult or impossible, so these boundaries must keep a test from producing any unintended or unauthorized fund movement.

Signing and withdrawal workflows require controlled scenarios that validate how the institution authenticates a request, applies policy, routes approvals, and reconciles the result. The RoE identifies designated test accounts, authorized participants, expected policy behavior, an upper limit for the scenario, and the evidence needed to validate the workflow. It then records the furthest permitted workflow step: request creation, policy decision, approval display, signing request, release decision, or reconciliation. Test identities are provisioned, used, and retired through the agreed access process. The same discipline applies to privileged access, customer data, and ledger operations: the access model, expected system behavior, and evidence boundary are agreed before testing begins; customer data is accessed only where necessary, minimized and redacted in evidence, and kept within the approved evidence repository.

These controls make it possible to test sensitive paths in production without treating them as ordinary application functions. They also give operations and the test team a common basis for coordinating live activity.

Testing and coordination

Testing and coordination are the live operating model for the engagement. They connect the service owner, security operations team, incident-response function, and test lead while the test is under way.

This model is necessary because expected test activity and a genuine security event can look similar. If monitoring, notifications, or pause decisions are unclear, testing can delay incident response or create uncertainty about whether production action is required.

The communications model identifies who knows the full test plan, who receives time-sensitive notices, and who can pause or resume activity. It varies by objective: some tests require close security operations center (SOC) coordination around sensitive systems, while limited-disclosure tests assess whether monitoring detects activity and escalation reaches the right people. Where a controlled scenario exercises a transaction-related workflow, the plan also covers relevant notifications and expected alerts from custody, wallet, transaction-screening, and monitoring providers. A separate safety contact and named pause authority remain reachable at all times. The service owner and test team agree measurable stop criteria, such as deviation from the error budget, unexpected increases in 95th-percentile (p95) response time or queue depth, suspicious account or wallet activity, material reconciliation discrepancies, or an unplanned security alert.

This preparation also strengthens operational resilience. The Hong Kong Securities and Futures Commission (SFC) expects virtual asset trading platform operators to maintain 24/7 monitoring and documented escalation procedures, and to conduct emergency and business-continuity drills with relevant third parties [3]. These regulatory references are illustrative, not legal advice; applicability is jurisdiction-specific and should be confirmed with counsel.

Figure 2. Testing and coordination control loop.
Figure 2. Testing and coordination control loop.

The diagram shows the control loop that applies while testing is under way. Its central point is that RoE boundaries do not end at authorization: monitoring and the safety channel turn those boundaries into decisions to resume, pause, contain, or hand a verified finding into remediation. It appears here because these decisions belong to live coordination, rather than to the earlier definition of scope or production preparation.

The same model governs critical findings: a demonstrated path to unauthorized fund movement, compromise of signing authority or a production control plane, exposure of highly sensitive customer data, or material risk to a critical service. The response path should identify the escalation threshold, the people who classify the finding, the authority to pause the relevant activity, and the process for containment, remediation, and validation. The penetration-testing team demonstrates and reports the issue; the institution retains authority for production decisions, customer communications, and remediation. A material-finding notification should use the protected safety channel first, followed by an agreed written record that does not expose unnecessary exploit detail.

Once a finding is contained and assigned, the value of the engagement depends on whether the institution can turn that result into a verified control improvement.

Remediation, retesting, and return to normal

Remediation and retesting are the closure process that converts a demonstrated weakness into a validated improvement. Returning to normal is part of the same process: test access, controlled scenarios, and collected evidence must not become new long-lived risk.

This final stage determines whether the engagement reduces risk or only produces a report. A finding without an owner, remediation path, and validation condition can remain open while the same attack path persists in production.

Before testing begins, identify how findings enter engineering, cloud, wallet-operations, or business-control workflows; who owns remediation; and which findings require a retest. At close, temporary test identities, permissions, and controlled scenarios return to their intended configuration, and service owners confirm that relevant health indicators remain within their expected range. The final record links each finding to its evidence, impact, owner, remediation plan, and retest condition. For a controlled transaction-related scenario, it also links the workflow reference, test identity, timestamp, approval record, and resulting ledger entry; any temporary access, approval, or test configuration created for the scenario is retired. Evidence is retained only for the agreed period, then securely destroyed or returned.

The result is not a vulnerability list but a tested, remediated, and retested set of controls that can support the institution's next operating decision.

Conclusion

Preparing for blockchain penetration testing is an institutional task. The institution defines the objective, operating context, authority, service constraints, and response model; the test team applies adversarial validation to that prepared environment.

With those elements in place, a penetration-testing engagement becomes more than a technical exercise. It becomes a controlled way to understand how the institution's deployed systems, people, and processes hold up under attack.

BlockSec helps institutions prepare and run that process: define the scope, map the operating environment, establish the RoE and production safeguards, and turn findings into remediated and retested controls. To plan the rules of engagement and production-safety controls for your next test, request a scoping conversation; more information is available upon request.

Continue with the series:

Also in this series, publishing soon:

  • Part 5: Authorization and Signing Security: Web, dApps, and
  • Part 6: Cloud and CI/CD Security: Automated Operations Attack Surfaces
  • Part 7: Treasury Control-Plane Security: Signing and Withdrawal Approvals
  • Part 8: Exchange Ledger Security: Theft Paths and Data-Plane Logic

The Blockchain Penetration Testing pillar page provides the engagement-level view.

Best Security Auditor for Web3

Validate design, code, and business logic before launch

References

Numbered in order of first appearance.

  1. National Institute of Standards and Technology, Rules of Engagement (ROE), CSRC Glossary.
  2. European Union, Regulation (EU) 2022/2554 — Digital Operational Resilience Act (DORA), Article 26.
  3. Hong Kong Securities and Futures Commission, Circular to Licensed Virtual Asset Trading Platform Operators on Custody of Virtual Assets (15 August 2025).

Best Security Auditor for Web3

Validate design, code, and business logic before launch. Aligned with the highest industry security standards.

BlockSec Audit
Rules of Engagement and Production Safety for Institutional Blockchain Penetration Testing