Skip to main content

Automation and Agents

Cursor Security Review: How to Triage AI Findings Safely

Cursor Security Review reports vulnerabilities on pull requests. Learn how to verify a finding, test a proposed fix, and keep human security ownership clear.

Editorial illustration of a human reviewing an AI security finding against code and evidence before making a decision. Official Cursor horizontal lockup.
On this page
  1. Treat the report as a hypothesis with a location
  2. Use a six-step triage loop
  3. 1. Locate the exact change
  4. 2. Trace the security-relevant flow
  5. 3. Check preconditions and impact
  6. 4. Seek evidence that could disprove the finding
  7. 5. Inspect the proposed fix independently
  8. 6. Record a scoped disposition
  9. Follow one illustrative authorization finding
  10. Use categories to guide questions, not to claim complete coverage
  11. Keep existing scanners and review ownership
  12. Tune team rules without hiding risk
  13. Evaluate the workflow with evidence you can defend
  14. A safe operating checklist
  15. Frequently asked questions
  16. Does Cursor Security Review replace SAST or manual code review?
  17. What should I do when a finding includes a proposed fix?
  18. Can I treat a dismissed finding as a permanent exception?
  19. Are Cursor's reported review improvements independently verified?

Cursor Security Review is Cursor’s pull-request security reviewer. Treat each finding as a hypothesis, not a disposition. A comment can point to the right line and still be wrong about exploitability, impact, or the fix. Test it against the exact code path, trust boundary, and attacker conditions. The tool can focus attention; a reviewer with system context must decide what the evidence supports.

Cursor's September 23, 2026 changelog describes Security Review as reading a pull request in the context of the codebase and posting one review comment with exploitable bugs, severity, attack path and proposed fix. It says draft pull requests are skipped, and places style and general code quality with Bugbot. These are Cursor's descriptions of the product. The sources reviewed for this article do not independently measure detection quality, false-positive rate, missed vulnerabilities, or time saved. Cursor's launch post includes reported efficiency and comment-acceptance figures, but it does not provide enough methodology in the material reviewed to treat them as a buyer's expected result.

Independent security guidance reinforces a narrower lesson: automated tools can help identify areas to investigate, but secure code review also depends on human understanding of application logic, data flow, changed trust boundaries and business requirements. OWASP's Secure Code Review Cheat Sheet explicitly describes manual review as complementary to automated security testing. It does not evaluate Cursor. The useful question is not whether to trust the bot in the abstract. Ask what evidence supports accepting, dismissing, or escalating this finding.

Break a finding into evidence. Severity is one part of a review, not the conclusion. CLAIM. What weakness is alleged?. LOCATION. Which changed path matters?. PATH. How does input or authority flow?. CONDITIONS. What must be true to exploit it?. REPAIR. What change closes the path?. View image detail

Choose Actual size to read the graphic closely.

Treat the report as a hypothesis with a location

A useful review starts by separating five things that a summary comment can make look like one:

  1. The claim: what weakness the reviewer believes exists.
  2. The location: the changed code, function, route, dependency, or configuration it points to.
  3. The path: how input, identity, permissions, or data moves to the risky operation.
  4. The conditions: what an attacker or user would need to do for the issue to matter.
  5. The proposed repair: the change that may address the claim, and what behavior it might alter.

Cursor says its finding includes severity, an attack path and proposed fix. That structure can help a human begin a focused review. It does not prove the path exists in the deployed system, that the described data is attacker-controlled, that the vulnerability is exploitable under your configuration, or that the proposed change closes the path. Each point needs to be checked against the actual code and system context.

This approach also prevents severity labels from carrying too much weight. “Critical” is not a substitute for impact analysis; “low” is not a reason to skip a plausible authorization defect. The label helps sort attention, while the review explains what the issue would let someone do, what must be true for that to happen, and what control already exists. A good disposition includes that reasoning, even if it is brief.

The evidence may vary by issue. A test can reproduce an input-validation failure. A call graph may reveal that a request handler applies an authorization check before data access. A configuration review may show that a route is only reachable behind a trusted boundary. A scanner may corroborate a known vulnerable dependency. For a business-logic issue, a security engineer may need to trace role changes, account ownership, or state transitions that ordinary static checks cannot infer reliably.

Do not ask for certainty that the evidence cannot supply. A test suite that does not cover a path is not proof that the path is safe. A static scanner that reports no issue is not proof that no exploit exists. A successful reproduction on one configuration confirms that scenario, not every tenant, identity provider, or deployment setting. State the scope of the evidence instead of converting one observation into a universal answer.

A six-step human triage loop. Each step can confirm, narrow or disprove the report. LOCATE. Verify current commit and diff. TRACE. Follow input, identity and data. ASSESS. Check preconditions and impact. COUNTERCHECK. Look for controls and disproof. TEST. Inspect patch and regression test. DISPOSITION. Record scope, owner and reason. View image detail

Choose Actual size to read the graphic closely.

Use a six-step triage loop

1. Locate the exact change

Open the affected lines and read the surrounding diff. Identify what changed and which behavior is new or different. A security reviewer can reason across the codebase, but the human still needs to anchor the report to a commit and path. Confirm that the comment points to the current PR version rather than a stale revision after new commits arrived.

If the finding names an existing helper, policy, or route, inspect that code too. A local line may appear unsafe without context, or a surrounding change may weaken a check elsewhere. Do not limit the review to the single highlighted line when the alleged problem depends on a caller, a shared authorization layer, a schema, or a runtime setting.

2. Trace the security-relevant flow

Follow untrusted input from its entry point to the sensitive operation. For authorization findings, trace identity and permission checks from request parsing to the resource access or state change. For an injection claim, follow the value through encoding, parameterization, validation and eventual interpreter. For secrets, determine whether the material is real, active and exposed beyond the intended boundary. For a dependency issue, confirm the exact package and version, affected component and reachable code path.

The important question is not whether a suspicious string exists near a sensitive function. It is whether data or authority crosses a boundary without the expected control. OWASP's guidance emphasizes changed components, new attack vectors, existing security controls, trust boundaries, authentication and authorization, data flow and business logic. Those are useful investigation targets for any reviewer, not proof that an AI system covers them completely.

3. Check preconditions and impact

Write down what must be true for exploitation. Can an unauthenticated user call the route? Does the attacker control the value? Is the action scoped to the current account? Is an upstream validation step mandatory or optional? Is the path exposed in production? Would the result disclose data, alter another user's records, execute a command, or only cause an internal warning?

This step narrows severity to the actual threat model. A security finding can be technically plausible but unreachable in the current system, or it can describe a reachable flaw whose impact is more serious than the label suggests. Conversely, an internal-only route can still present risk if untrusted jobs can reach it or if a trusted boundary is weak. The conclusion must explain why the preconditions hold or do not hold.

4. Seek evidence that could disprove the finding

Reviewers naturally look for confirming clues after a tool flags an issue. Deliberately look for counter-evidence too. Is there a server-side check on the complete request path? Does a framework enforce the constraint? Is the input already canonicalized? Does a later transaction guard prevent the unsafe state? Is the code path excluded in the deployed configuration? Can a regression test capture the behavior?

The goal is not to dismiss the finding. It is to test its strongest version. If the claim is “this user can read another account's invoice,” try to establish the principal, target record, policy evaluation, and data access sequence. If the evidence is incomplete, record that and ask for a specialist review rather than assuming the code is safe. A reasoned “needs security owner” is better than a premature close.

5. Inspect the proposed fix independently

A repair that blocks one path can break legitimate behavior or leave a parallel path open. Check that the change enforces the intended control at the appropriate boundary, uses the correct identity or tenant context, and cannot be bypassed by alternate callers. Read the new tests and verify they would fail on the original behavior. Run the repository's relevant checks, then review the test output and code diff rather than accepting “tests pass” as a generic guarantee.

If an agent generates the patch, keep the same code review expectations that apply to any contributor. Cursor's description of a one-click fix says how the interface presents remediation; it does not establish that the repair has been validated in your service. Treat it as suggested code. Inspect its assumptions, the blast radius and the regression path before merge.

6. Record a scoped disposition

Use a small vocabulary so future readers can distinguish outcomes. For example: confirmed and fixed, confirmed and accepted temporarily with owner/date, not exploitable under documented conditions, duplicate of tracked finding, insufficient evidence, or needs specialist review. Attach the commit reviewed, evidence or test, affected path, reason, owner and follow-up date where relevant.

“Dismissed” by itself is too thin. A later code change, configuration change or different tenant model can invalidate the original reason. Record why this PR's claim did or did not hold. This helps avoid rediscovering the same issue without teaching future reviewers that a broad class of findings is harmless.

Authorization example: tenant boundary. Hypothetical flow for review, not a reported Cursor finding. WORKSPACE A. Manager identity. Requests project A. REQUEST. User-controlled project ID. Do not trust the client alone. SERVER CHECK. Tenant-scoped authorization. Check before enqueueing work. WORKSPACE B. Different tenant record. Expected outcome: denied. View image detail

Choose Actual size to read the graphic closely.

Follow one illustrative authorization finding

Imagine a pull request adds an endpoint that allows a manager to export project records. A reviewer reports a possible authorization bypass: the handler loads a project by an ID from the request and starts an export, but the diff does not show a check that the manager belongs to that project's workspace.

The first question is whether the report matches the code at the current commit. The reviewer opens the endpoint and follows the call into the export service. Then they trace where the authenticated user and workspace membership are checked. Perhaps the project loader applies a tenant constraint for every caller. Perhaps the UI hides cross-workspace projects, but the endpoint itself does not enforce the restriction. Those produce very different outcomes.

Next, the reviewer checks the preconditions. Can a manager choose an arbitrary project identifier, or does the server derive it from a scoped session? Is the route behind a separate authorization middleware? Does that middleware cover every request path? The team may add a test with a manager from workspace A requesting a project from workspace B. A passing test should prove a denial for that boundary and continue to prove an allowed export for a manager in workspace A.

Now inspect the proposed fix. Does it use the authenticated tenant, or does it compare a caller-supplied workspace value with another caller-supplied project ID? Does it run before the export job is enqueued? Could the asynchronous worker later use a broader service identity without carrying the original authorization decision? A one-line check at the controller might not be sufficient if the background job can be invoked through another path.

Suppose the existing service already applies an authoritative workspace filter, and tests cover the cross-tenant case. The reviewer might close the finding as not exploitable in this path, citing the service check and test. If the control exists only in a frontend screen, the issue is likely real and needs repair. If the worker's access model remains unclear, the right disposition is “insufficient evidence” or “needs specialist review,” not “false positive.” This example is hypothetical. It does not describe a tested Cursor finding or a Rise security assessment.

Keep each control in its lane. Different tools produce different evidence. SAST. Known patterns and code paths. Not a business-logic verdict. SCA. Package identity and advisories. Not proof of reachability. AI REVIEW. Contextual questions and leads. Not an independent sign-off. HUMAN + TESTS. Threat model and behavior evidence. Still bounded by test coverage. View image detail

Choose Actual size to read the graphic closely.

Use categories to guide questions, not to claim complete coverage

Cursor's changelog says Security Review looks for injection across SQL, command and template surfaces; authentication and authorization bypasses; secrets and credentials; server-side request forgery and unvalidated redirects; unsafe deserialization; and dependency changes that introduce known vulnerabilities. It also says the reviewer traces input through the code and can apply team rules. Those categories are useful starting points. They are not a claim that every vulnerability in those classes will be found or every rule will be interpreted correctly.

For injection, ask where the input comes from, what interpreter receives it, and whether the framework's parameterization or escaping applies to this exact path. For authorization, check the subject, action, object and tenant boundary. For secrets, validate whether the value is a secret rather than test data and where it has propagated. For SSRF, map allowed destinations, redirects, DNS behavior and network boundaries. For unsafe deserialization, identify the actual format and attacker control. For a vulnerable dependency, verify package identity, version, reachability and the vendor advisory's affected range.

These questions also expose different tool roles. A dependency scanner can match package versions to advisories. A secret scanner can search known patterns. A static analyzer can find code patterns. An AI reviewer may add context about data flow or the purpose of a change. A human can reason about business authorization and the system's real threat model. The overlap is useful, but one tool's output does not replace every other source of evidence.

The same OWASP guidance describes manual secure code review as complementary to automated security testing, particularly for application logic and context-specific vulnerabilities. NIST's Secure Software Development Framework likewise places security practices inside the software development lifecycle rather than treating one control as the whole program. Neither source evaluates Cursor Security Review. They support the operating principle that review tools belong inside a broader development and verification process.

From severity to a defensible decision. A label routes attention; evidence justifies disposition. Reachability. Caller, route, data path. Which deployment/config?. Impact. Action, object, identity. What can an attacker do?. Control. Policy and call chain. Where is enforcement?. Repair. Patch + failing test. What behavior now changes?. View image detail

Choose Actual size to read the graphic closely.

Keep existing scanners and review ownership

An AI reviewer should fit into the system your team already uses. Keep SAST, dependency scanning, secret detection, tests, threat modeling and human review where they add distinct evidence. If the same check runs twice, decide whether the second result adds context or just noise. If there is an overlap, establish which system is authoritative for a blocking rule and how disagreements are resolved.

Do not disable a deterministic scanner just because an AI agent can discuss the same category. A scanner may provide reproducible results on a known rule or dependency database. An AI reviewer may identify a contextual question the scanner did not surface. They fail in different ways. The strongest operating model preserves the evidence from each, avoids double-counting the same issue as two independent confirmations, and assigns an owner to reconcile disagreement.

GitHub's Code Scanning documentation provides a useful baseline for what an established finding workflow may expose: the affected line, severity, issue type, and pull-request annotations, with additional data-flow details for some CodeQL findings. Compare the evidence and ownership path your team already gets from those tools before adding another comment stream. This is workflow context, not a comparison test of Cursor Security Review. GitHub Docs: Code scanning alerts

The review needs an owner with the right expertise. A pull-request author can correct many straightforward issues, but sensitive authorization, cryptographic, tenant-isolation or data-exposure questions may need a security engineer or system owner. Define which findings require escalation and what evidence the reviewer must attach. A severity label can help route attention; the actual code path and possible impact decide who needs to review it.

Cursor's September changelog says Security Review is available on Teams and Enterprise plans. That does not confirm access, billing or usage controls for a particular account. Before activation, have a team admin check the current dashboard and contract, then choose the repositories to review. Cursor says draft pull requests are skipped, so include that boundary in the team's review policy.

Treat custom rules as policy code. Tune scope without suppressing the risk category. DEFINE. Owner, path and invariant. EXAMPLE. Positive and safe negative case. REVIEW. Check findings and misses. UPDATE. Version and reassess exceptions. View image detail

Choose Actual size to read the graphic closely.

Tune team rules without hiding risk

Cursor's changelog says teams can add codebase-specific rules, such as requiring external calls through a particular client or restricting which tables are queried from a request handler. These rules can express local architecture choices that generic guidance may not know. A rule is only useful if it is clear, scoped to the right paths, maintained as the architecture changes, and tested against representative examples.

Avoid writing a rule that tells the agent to suppress an entire category because past findings were noisy. If SQL injection reports often lack enough context, add the relevant parameterization pattern and tell reviewers what evidence to inspect. If a trusted wrapper consistently enforces authorization, name the wrapper and define how changes to it should be escalated. If findings around test fixtures are irrelevant, scope that path without excluding production code that resembles it.

Treat the rule set like policy code. Keep an owner, change history and examples that include both an expected finding and a safe negative case. When changing a rule, review a sample of prior findings if possible so that reduced noise does not come from missed coverage. A dropped alert count alone is ambiguous: fewer findings could mean clearer rules, or it could mean the system stopped surfacing important cases.

The same caution applies to dismissing findings. Cursor says dismissing a finding with a reason prevents the reviewer from raising it again on that PR. That may help a single PR avoid repeated comments, but it does not prove that a similar issue elsewhere is safe. Capture the reason on the specific change and preserve the underlying system rule or exception in the team's normal security process when it applies more broadly.

Evaluate more than posted comments. Include quiet PRs so missed issues stay visible. CONFIRMED. Evidence + affected class. REJECTED. Reason and counter-evidence. UNRESOLVED. Owner and next review. NO FINDING. Qualified sample review. View image detail

Choose Actual size to read the graphic closely.

Evaluate the workflow with evidence you can defend

Cursor's launch post reports changes in average review time and comment acceptance. The reported average review time moves from 4.8 to 3.8 minutes, and acceptance from a 45 to 50 percent range to a 60 to 70 percent range. The page does not provide, in the reviewed material, the sample construction, task mix, confidence range or independent replication needed to project the same result onto another team's codebase. Attribute the figures as Cursor-reported and do not use them as a target or ROI estimate without more detail.

A team's pilot should measure whether the workflow improves review quality and decision clarity. Begin with a defined set of pull requests and preserve the existing safeguards. For each finding, record whether it was confirmed, a false alarm, already known, or unresolved; the affected class; what evidence was needed; whether the proposed fix was accepted, changed, or rejected; and whether tests captured the case. Review a sample of PRs with no findings too, because missed issues are invisible if you only count posted comments.

Be careful with “precision.” If you count only posted findings, you may not know what the system missed. For a meaningful evaluation, a qualified reviewer needs a defined reference set or independent review of at least a sample of changes. A high acceptance rate can mean useful findings, but it can also reflect reviewer trust, convenience or pressure to clear comments. The metric needs an operational definition and evidence about rejected and missed cases.

Likewise, shorter review time is not automatically safer or more productive. If a reviewer spends less time because the findings are trivial, that may be valuable. If they spend less time because they trust an unverified suggestion, that may hide risk. Compare the saved effort with the depth of reasoning, test quality, rework and incident outcomes. Do not claim causal improvement from a small pilot without describing its scope and limitations.

The most useful initial measures may be modest: Are findings anchored to the correct commit? Can reviewers explain why an issue is confirmed or not? Do the generated fixes include tests for the actual risk? Which categories require specialist escalation? How often do changes in team rules alter the result? These measures will not prove comprehensive security, but they help a team understand whether its specific workflow is becoming more legible.

Choose a scoped disposition. A reasoned close is more useful than “dismissed.” Confirmed and fixed, with test evidence. Confirmed; temporary exception has owner/date. Not exploitable in this reviewed path. Duplicate of a tracked issue. Insufficient evidence; keep open. Specialist review needed before merge. View image detail

Choose Actual size to read the graphic closely.

A safe operating checklist

Before enabling Cursor Security Review on repositories, decide:

  • Which repositories and pull-request types are in scope, and that draft PRs are currently skipped.
  • Which security tools remain required and which control is authoritative for each category.
  • Who triages findings by severity and vulnerability class.
  • Which issues require a security specialist or service owner.
  • How reviewers confirm the exact commit, code path, preconditions and impact.
  • What evidence supports a fix, a dismissal or an unresolved status.
  • How test changes demonstrate the control, including negative and allowed cases.
  • Where team-specific rules live, who maintains them and how exceptions are reviewed.
  • Whether account plan, usage pool and admin access match the intended rollout.
  • How the team samples PRs with no AI findings to look for missed issues.

If the organization cannot name a person who owns the final security decision, enabling a reviewer does not fill that governance gap. Start by assigning responsibility and making the evidence path usable. Then a tool's finding can be one input in a clear process instead of a new comment people either accept automatically or learn to ignore.

The related Rise guide, Claude Opus 5.5 for Unattended Coding: Where Should Review Stay?, addresses a neighboring decision about review boundaries for coding agents. It is not product documentation for Cursor, but it can help a team keep the question of authority separate from the question of whether a particular finding is correct. Teams planning broader security controls may also find Base44 Base Code Security: GitHub Access and Shared Secrets useful as a distinct integration-control example.

Frequently asked questions

Does Cursor Security Review replace SAST or manual code review?

No. Cursor describes a PR security reviewer, while OWASP says automated tools complement manual secure code review. Keep the controls that address different evidence and retain human responsibility for context and disposition.

What should I do when a finding includes a proposed fix?

Review the affected path and the fix independently. Verify that it enforces the intended control, does not create another bypass, and includes tests that would fail on the original behavior. A proposed patch is not proof that the vulnerability is resolved.

Can I treat a dismissed finding as a permanent exception?

Not automatically. Record why the report did not apply to the reviewed PR and whether that reason generalizes. A changed code path or architecture can invalidate an earlier dismissal.

Are Cursor's reported review improvements independently verified?

Not by the sources examined for this package. Cursor's launch post reports the figures; the inspected source does not supply enough methodology for a transferable estimate. A team should evaluate its own workflow with a defined review sample and include PRs with no findings.

This guide focuses on deciding what a security finding proves. For the separate question of monitoring a reviewed change after it ships, see Cursor Rollouts: How to Monitor a Pull Request Through Production. A runtime monitor does not substitute for code review, and a code finding does not establish production behavior.

Checked for this article

Sources

  1. Cursor, Rollouts and Security Review changelog
  2. Cursor, Bots for the last mile: Rollouts, Security Review
  3. OWASP Cheat Sheet Series, Secure Code Review
  4. NIST SP 800-218, Secure Software Development Framework
  5. GitHub documentation, Code scanning alerts

Keep going

All articles