CWE-20: Improper Input Validation

What is CWE-20?

The product accepts input or data without correctly verifying that it has the properties required for safe and correct processing.

Data statistics

OWASP TOP 10:2025 RANK5 — A05:2025 — Injection
RELATED CVES (365 DAYS)869
ABSTRACTIONClass
LIKELIHOOD OF EXPLOITHigh

Vulnerabilities mapped to CWE-20

869 vulnerabilities262.1% increase year over year

Vulnerabilities in CISA KEV for CWE-20

8 vulnerabilities100% increase year over year

Official definition

ByMitre CWE

The product receives input or data, but it does not validate or incorrectly validates that the input has the properties that are required to process the data safely and correctly.

Input validation is a frequently-used technique for checking potentially dangerous inputs in order to ensure that the inputs are safe for processing within the code, or when communicating with other components.

Input can consist of:

  • raw data - strings, numbers, parameters, file contents, etc.

  • metadata - information about the raw data, such as headers or size

Data can be simple or structured. Structured data can be composed of many nested layers, composed of combinations of metadata and raw data, with other simple or structured data.

Many properties of raw data or metadata may need to be validated upon entry into the code, such as:

  • specified quantities such as size, length, frequency, price, rate, number of operations, time, etc.

  • implied or derived quantities, such as the actual size of a file instead of a specified size

  • indexes, offsets, or positions into more complex data structures

  • symbolic keys or other elements into hash tables, associative arrays, etc.

  • well-formedness, i.e. syntactic correctness - compliance with expected syntax

  • lexical token correctness - compliance with rules for what is treated as a token

  • specified or derived type - the actual type of the input (or what the input appears to be)

  • consistency - between individual data elements, between raw data and metadata, between references, etc.

  • conformance to domain-specific rules, e.g. business logic

  • equivalence - ensuring that equivalent inputs are treated the same

  • authenticity, ownership, or other attestations about the input, e.g. a cryptographic signature to prove the source of the data

Implied or derived properties of data must often be calculated or inferred by the code itself. Errors in deriving properties may be considered a contributing factor to improper input validation.

Detailed description

Input validation applies to raw data and metadata, including strings, numbers, parameters, file contents, headers, and sizes. Validation may need to establish size and range limits, indexes and offsets, symbolic keys, syntax and lexical correctness, data type, consistency between related values, conformance to business rules, equivalence, and authenticity or ownership claims.

The data may be simple or deeply structured, with nested combinations of metadata and raw values. Properties that are implied or derived by the application must also be calculated correctly. Validation failures can occur when developers trust client-side checks, cookies, hidden fields, or other inputs that an attacker can modify, or when validation is incomplete, occurs before data from multiple sources is combined, or is bypassed through decoding and representation differences.

Characteristics

This weakness is a broad validation failure rather than a problem limited to one input format. It can affect network data, request parameters and headers, cookies, environment variables, reverse DNS results, URL components, email, files and filenames, databases, API results, and data supplied by external systems. The failure may involve missing values, extra values, malformed syntax, invalid types, out-of-range quantities, inconsistent length or size fields, unsafe references, or violations of domain-specific rules.

It is commonly introduced during architecture and design when trust boundaries are misunderstood, and during implementation when checks are omitted or applied inconsistently. It is especially important at component interfaces, language boundaries, and transitions between external representations and internal data.

Common consequences

Depending on the input and the affected component, an attacker may cause crashes, restarts, excessive CPU or memory consumption, or other denial of service. Control over resource references may expose memory, files, or directories. Malicious input may also modify memory or data, alter control flow, or result in unauthorized code or command execution.

ImpactScopeExplanation
DoS: Crash, Exit, or Restart, DoS: Resource Consumption (CPU), DoS: Resource Consumption (Memory)AvailabilityAn attacker could provide unexpected values and cause a program crash or arbitrary control of resource allocation, leading to excessive consumption of resources such as memory and CPU.
Read Memory, Read Files or DirectoriesConfidentialityAn attacker could read confidential data if they are able to control resource references.
Modify Memory, Execute Unauthorized Code or CommandsIntegrity, Confidentiality, AvailabilityAn attacker could use malicious input to modify data or possibly alter control flow in unexpected ways, including arbitrary command execution.

Risk mitigations

Architecture and design:

  • Reduce the attack surface by identifying every path through which untrusted data can enter, including indirect inputs obtained through API calls.
  • Consider language-theoretic security techniques. Use a formal input language and a distinct parsing and recognition layer that separates raw input from internal representations.
  • Use established validation libraries or frameworks such as Struts or the OWASP ESAPI Validation API, while reviewing their use because a framework does not eliminate validation errors.
  • Duplicate security checks on the server when checks also run on the client. Client-side checks remain useful for feedback, intrusion indicators, and reducing accidental errors, but cannot protect the server by themselves.

Implementation:

  • Assume all input may be malicious and prefer an allowlist, or “accept known good,” strategy. Check length, type, full value ranges, missing and extra fields, syntax, cross-field consistency, and business rules. Denylists may supplement detection or rejection of obviously malformed input, but should not be the sole defense.
  • Validate data after combining values from multiple sources, because individually valid elements may violate restrictions when combined.
  • Validate carefully across language boundaries, including calls from interpreted code to native code.
  • Convert input directly to the expected type, then check its permitted range and consistency with related fields.
  • Decode and canonicalize input into the application's internal representation before validation, avoid unintended repeated decoding, and use canonicalization controls such as the OWASP ESAPI Canonicalization control.
  • Use the same character encoding between communicating components and explicitly set the encoding where the protocol permits it.
  1. Attack Surface Reduction · Architecture and DesignConsider using language-theoretic security (LangSec) techniques that characterize inputs using a formal language and build "recognizers" for that language. This effectively requires parsing to be a distinct layer that effectively enforces a boundary between raw input and internal data representations, instead of allowing parser code to be scattered throughout the program, where it could be subject to errors or inconsistencies that create weaknesses. [REF-1109] [REF-1110] [REF-1111]
  2. Libraries or Frameworks · Architecture and DesignUse an input validation framework such as Struts or the OWASP ESAPI Validation API. Note that using a framework does not automatically address all input validation problems; be mindful of weaknesses that could arise from misusing the framework itself (CWE-1173).
  3. Attack Surface Reduction · Architecture and Design, ImplementationUnderstand all the potential areas where untrusted inputs can enter the product, including but not limited to: parameters or arguments, cookies, anything read from the network, environment variables, reverse DNS lookups, query results, request headers, URL components, e-mail, files, filenames, databases, and any external systems that provide data to the application. Remember that such inputs may be obtained indirectly through API calls.
  4. Input Validation · Implementation · Effectiveness: HighAssume all input is malicious. Use an "accept known good" input validation strategy, i.e., use a list of acceptable inputs that strictly conform to specifications. Reject any input that does not strictly conform to specifications, or transform it into something that does. When performing input validation, consider all potentially relevant properties, including length, type of input, the full range of acceptable values, missing or extra inputs, syntax, consistency across related fields, and conformance to business rules. As an example of business rule logic, "boat" may be syntactically valid because it only contains alphanumeric characters, but it is not valid if the input is only expected to contain colors such as "red" or "blue." Do not rely exclusively on looking for malicious or malformed inputs. This is likely to miss at least one undesirable input, especially if the code's environment changes. This can give attackers enough room to bypass the intended validation. However, denylists can be useful for detecting potential attacks or determining which inputs are so malformed that they should be rejected outright.
  5. Architecture and DesignFor any security checks that are performed on the client side, ensure that these checks are duplicated on the server side, in order to avoid CWE-602. Attackers can bypass the client-side checks by modifying values after the checks have been performed, or by changing the client to remove the client-side checks entirely. Then, these modified values would be submitted to the server. Even though client-side checks provide minimal benefits with respect to server-side security, they are still useful. First, they can support intrusion detection. If the server receives input that should have been rejected by the client, then it may be an indication of an attack. Second, client-side error-checking can provide helpful feedback to the user about the expectations for valid input. Third, there may be a reduction in server-side processing time for accidental input errors, although this is typically a small savings.
  6. ImplementationWhen your application combines data from multiple sources, perform the validation after the sources have been combined. The individual data elements may pass the validation step but violate the intended restrictions after they have been combined.
  7. ImplementationBe especially careful to validate all input when invoking code that crosses language boundaries, such as from an interpreted language to native code. This could create an unexpected interaction between the language boundaries. Ensure that you are not violating any of the expectations of the language with which you are interfacing. For example, even though Java may not be susceptible to buffer overflows, providing a large argument in a call to native code might trigger an overflow.
  8. ImplementationDirectly convert your input type into the expected data type, such as using a conversion function that translates a string into a number. After converting to the expected data type, ensure that the input's values fall within the expected range of allowable values and that multi-field consistencies are maintained.
  9. ImplementationInputs should be decoded and canonicalized to the application's current internal representation before being validated (CWE-180, CWE-181). Make sure that your application does not inadvertently decode the same input twice (CWE-174). Such errors could be used to bypass allowlist schemes by introducing dangerous inputs after they have been checked. Use libraries such as the OWASP ESAPI Canonicalization control. Consider performing repeated canonicalization until your input does not change any more. This will avoid double-decoding and similar scenarios, but it might inadvertently modify inputs that are allowed to contain properly-encoded dangerous content.
  10. ImplementationWhen exchanging data between components, ensure that both components are using the same character encoding. Ensure that the proper encoding is applied at each interface. Explicitly set the encoding you are using whenever the protocol allows you to do so.

Detection methods

Use multiple complementary approaches because no single method covers all validation rules. Automated static analysis can identify locations where recognized validation methods or frameworks are absent, but may produce false positives when custom validation is not understood. Manual static analysis is necessary for custom rules, especially business logic.

Fuzzing should supply unexpected inputs and verify that the application remains stable and returns application-controlled errors rather than crashes, exceptions, or interpreter-generated messages. The record also identifies binary or bytecode disassembly and weakness analysis, web, web-service, and database scanners, fuzz testers and framework-based fuzzers, host interface scanning, monitored virtual environments, focused source spot checks, manual source review, source-code weakness analyzers, inspections, formal methods, and attack modeling. The SOAR-listed binary, bytecode, and partial-coverage methods are not complete coverage guarantees.

MethodApproachEffectiveness
Automated Static AnalysisSome instances of improper input validation can be detected using automated static analysis. A static analysis tool might allow the user to specify which application-specific methods or functions perform input validation; the tool might also have built-in knowledge of validation frameworks such as Struts. The tool may then suppress or de-prioritize any associated warnings. This allows the analyst to focus on areas of the software in which input validation does not appear to be present. Except in the cases described in the previous paragraph, automated static analysis might not be able to recognize when proper input validation is being performed, leading to false positives - i.e., warnings that do not have any security consequences or require any code changes.—
Manual Static AnalysisWhen custom input validation is required, such as when enforcing business rules, manual analysis is necessary to ensure that the validation is properly implemented.—
FuzzingFuzzing techniques can be useful for detecting input validation errors. When unexpected inputs are provided to the software, the software should not crash or otherwise become unstable, and it should generate application-controlled error messages. If exceptions or interpreter-generated error messages occur, this indicates that the input was not detected and handled within the application logic itself.—
Automated Static Analysis - Binary or BytecodeAccording to SOAR [REF-1479], the following detection techniques may be useful: ``` Cost effective for partial coverage: ``` Bytecode Weakness Analysis - including disassembler + source code weakness analysis Binary Weakness Analysis - including disassembler + source code weakness analysisSOAR Partial
Manual Static Analysis - Binary or BytecodeAccording to SOAR [REF-1479], the following detection techniques may be useful: ``` Cost effective for partial coverage: ``` Binary / Bytecode disassembler - then use manual analysis for vulnerabilities & anomaliesSOAR Partial
Dynamic Analysis with Automated Results InterpretationAccording to SOAR [REF-1479], the following detection techniques may be useful: ``` Highly cost effective: ``` Web Application Scanner Web Services Scanner Database ScannersHigh
Dynamic Analysis with Manual Results InterpretationAccording to SOAR [REF-1479], the following detection techniques may be useful: ``` Highly cost effective: ``` Fuzz Tester Framework-based Fuzzer ``` Cost effective for partial coverage: ``` Host Application Interface Scanner Monitored Virtual Environment - run potentially malicious code in sandbox / wrapper / virtual machine, see if it does anything suspiciousHigh
Manual Static Analysis - Source CodeAccording to SOAR [REF-1479], the following detection techniques may be useful: ``` Highly cost effective: ``` Focused Manual Spotcheck - Focused manual analysis of source Manual Source Code Review (not inspections)High
Automated Static Analysis - Source CodeAccording to SOAR [REF-1479], the following detection techniques may be useful: ``` Highly cost effective: ``` Source code Weakness Analyzer Context-configured Source Code Weakness AnalyzerHigh
Architecture or Design ReviewAccording to SOAR [REF-1479], the following detection techniques may be useful: ``` Highly cost effective: ``` Inspection (IEEE 1028 standard) (can apply to requirements, design, source code, etc.) Formal Methods / Correct-By-Construction ``` Cost effective for partial coverage: ``` Attack ModelingHigh

Representative vulnerabilities

The official record provides the following representative examples, not an exhaustive list. They include:

  • CVE-2024-37032: an unvalidated digest format enabled relative path traversal in an LLM management tool.
  • CVE-2022-45918: an improperly validated path enabled traversal using ../ sequences.
  • CVE-2021-30860 and CVE-2021-30663: improper validation led to integer overflow.
  • CVE-2021-22205: validation bypass through a backslash followed by a newline led to eval injection.
  • CVE-2021-21220: insufficient validation led to heap corruption.
  • CVE-2020-9054, CVE-2008-5305, and CVE-2008-1625: insufficient validation enabled command or eval injection and code execution.
  • CVE-2020-3452, CVE-2008-1284, and CVE-2008-3660: validation failures contributed to directory traversal.
  • CVE-2020-3580, CVE-2008-3843, and CVE-2008-2223: validation failures contributed to XSS or SQL injection.
  • Other listed examples involve malformed or inconsistent packet fields, missing parameters, zero-length values, invalid versions, infinite loops, crashes, information exposure, memory over-read or corruption, resource consumption, HTTP response smuggling, security bypass, and code execution.

Below are representative vulnerabilities related to this CWE, prioritized by severity.

Sources (13)

CWE™ Program, operated by The MITRE Corporation. Copyright © 2006–2026, The MITRE Corporation. The MITRE Corporation hereby grants you a non-exclusive, royalty-free license to use CWE for research, development, and commercial purposes. CWE Terms of Use.

Learn more

Run an in-depth assessment with complete web risk management

CyStack VulnScan continuously discovers assets, validates vulnerabilities, and helps security teams prioritize remediation across the organization.

Explore CyStack VulnScan