ISTQB Advanced (CTAL-TM v3.0) Mock Exam #2 — Questions & Answers
Every question in this mock exam, with the correct answer marked and a written rationale behind each option below — for reading and review, not a timed run.
Chapter 1 · Managing the Test Activities — 26 questions
During test planning for a national rail ticketing platform, the test manager…
During test planning for a national rail ticketing platform, the test manager uses the test policy and the organizational test strategy as inputs and identifies a product risk: seat double-booking under peak sales load. The team decides to perform a design review and static analysis of the reservation locking logic before that logic is implemented. Which risk treatment approach does this decision illustrate?
Preventive: the review and static analysis are carried out before the locking logic is implemented, so that the defect is stopped from being introduced in the first place.
Correct. Preventive treatment acts ahead of implementation to stop the risk from materializing, which is exactly what a design review and static analysis of not-yet-written logic achieve.
Corrective: the review and static analysis are carried out to find and repair defects that have already been introduced into the implemented locking logic.
Incorrect. Corrective treatment addresses defects that already exist in the test object; here the measure is scheduled before the logic is written.
Mitigating: the review and static analysis are carried out to reduce the business impact of a double-booking that has already reached ticket buyers, rather than to stop it occurring.
Incorrect. Mitigation reduces the consequences of a failure that occurs; the described measure is aimed at preventing the defect, not at limiting damage afterwards.
Accepting: the review and static analysis are recorded as evidence that the risk is tolerated, with production incidents monitored instead of any further treatment being planned.
Incorrect. Acceptance means no treatment action is taken beyond monitoring; the team here has actively scheduled test activities against the risk.
Test planning takes the test policy and the organizational test strategy as inputs, includes identification of product risks, and selects approaches to treat those risks. A measure applied before the code exists, so that the defect is never introduced, is a preventive approach. Corrective measures find and repair defects that already exist, mitigating measures reduce the impact of a failure, and acceptance means tolerating the risk without further treatment.
On a smart energy-metering programme, the test manager checks each week whether the conditions for starting system testing are satisfied, and later whether the agreed exit criteria have been met so that the test level can be approved as complete. Where do these two checks belong in the test process?
Both are test monitoring and control activities: measurements gathered during testing are compared with the planned targets, and control actions are taken where a deviation is found.
Correct. Checking test readiness and approving completion against exit criteria are monitoring activities, with control actions following from any deviation.
Both are test analysis activities: readiness and completeness are established by deriving test conditions from the test basis and evaluating the coverage those conditions give.
Incorrect. Test analysis identifies what to test by examining the test basis; it does not assess readiness to start or approval to finish a test level.
Test readiness is a test planning activity and approval against exit criteria is a test completion activity, so neither belongs to test monitoring and control, whose scope is progress reporting.
Incorrect. Planning defines the criteria and completion archives the outcome, but the act of checking status against those criteria during the project is monitoring and control.
Both are test implementation activities: the test manager assembles the testware and confirms that the test environment is ready before and after the agreed execution window.
Incorrect. Test implementation prepares the testware, data and environment needed to run tests; it does not evaluate progress against planned targets.
Checking test readiness and approving completion against exit criteria are both test monitoring and control activities. Monitoring collects measurements about the testing and compares them with the planned targets; control consists of the corrective actions taken when a deviation from the plan is found.
A logistics tracking company develops its embedded telematics firmware in a…
A logistics tracking company develops its embedded telematics firmware in a sequential lifecycle while its web and mobile fleet portal is developed using Scrum. Management describes the result as a deliberate hybrid development model. Which explanation matches the way hybrid models are characterized in the syllabus?
Hybrid models arise either from a transition to Agile or from a fit-for-purpose decision in which each part uses the lifecycle that suits it best; testers can coordinate across the parts using a scrum-of-scrums.
Correct. These are the two motivations given for hybrid models, and the scrum-of-scrums is the recommended coordination mechanism for testers spanning them.
Hybrid models arise from a transition to Agile, and the sequential parts are expected to be converted to Scrum before the next release; coordination during that transition is handled by a scrum-of-scrums.
Incorrect. Transition to Agile is only one of the two motivations, and a hybrid model may be a lasting fit-for-purpose choice rather than a stage on the way to a fully Agile lifecycle.
Hybrid models are defined by combining the test levels of a sequential lifecycle with the test types of an Agile lifecycle, and coordination between the two parts is handled in the joint release retrospective.
Incorrect. A hybrid model combines development lifecycles for different parts of the product, not test levels with test types, and retrospectives are not the coordination forum described.
Hybrid models require that testing for every part is planned in a single iteration-independent test plan, and the scrum-of-scrums is the forum in which the product owner formally approves that combined plan before each release begins.
Incorrect. Hybrid models do not mandate a single test plan, and the scrum-of-scrums is a coordination meeting rather than a plan approval gate.
Hybrid models usually arise from one of two motivations: an organization that is transitioning to Agile and still has sequential parts, or a fit-for-purpose choice where each part of the product uses the lifecycle best suited to it. Testers working across the parts can coordinate through a scrum-of-scrums.
A national tax filing portal is moving part of its development from a V-model to Scrum. The test manager is briefing the team on what will change. Considering only estimation and testware, which statement correctly contrasts the sequential and the iterative model?
In the sequential model, estimation is performed up front for the whole project and testware is detailed, formally maintained documentation; in the iterative model, estimation is performed by the team per iteration and testware is lighter and evolves with the product.
Correct. This is the contrast the syllabus draws for these two aspects: up-front whole-project estimation with formal testware versus per-iteration team estimation with lightweight, evolving testware.
In the sequential model, estimation is re-done by the whole team at the start of every phase and testware is deliberately kept lightweight; in the iterative model, estimation is fixed for the whole release and testware is maintained as formal specifications.
Incorrect. This reverses both aspects: the sequential model uses up-front estimation and formal testware, while the iterative model estimates per iteration and keeps testware lightweight.
In both models estimation is performed once per release by the test manager alone; the only difference is that sequential testware is reusable across later projects while iterative testware is discarded at the end of each iteration once the increment has been accepted at the review.
Incorrect. Estimation ownership and timing differ between the models, and iterative testware is retained and evolved rather than discarded after each iteration.
In the sequential model, estimation is based on story points agreed in planning poker and testware is held in the product backlog; in the iterative model, estimation is based on a work breakdown structure and testware is filed with the master test plan.
Incorrect. The estimation techniques and testware locations are swapped: story points and the backlog belong to the iterative model, the work breakdown structure and master test plan to the sequential one.
In a sequential model such as the V-model, test estimation is normally performed up front for the whole project and testware consists of detailed, formally maintained documentation. In an iterative model such as Scrum, estimation is performed by the team for each iteration and testware is lighter weight, evolving together with the product.
A test manager on a university admissions platform is preparing the product risk analysis for the upcoming clearing period, when application volumes peak. Which THREE of the following are risk identification techniques? (Select THREE)
Interviewing subject matter experts from admissions operations about what could go wrong during clearing.
Correct. Expert interviews are a listed risk identification technique and draw on knowledge that the test team may not hold.
Holding a retrospective on the previous clearing release to capture lessons learned about failures that occurred.
Correct. Retrospectives and lessons learned from earlier projects are a listed source for identifying product risks.
Running a risk workshop, for example using FMEA, with development, operations and admissions representatives.
Correct. Risk workshops, including techniques such as FMEA, are a listed risk identification technique.
Rating the likelihood and the impact of each identified item as low, medium or high in an agreed risk matrix.
Incorrect. Assigning likelihood and impact ratings is risk assessment, which takes place after the risks have been identified.
Selecting mitigation actions and allocating the available test effort to the highest-rated risk items first.
Incorrect. Choosing treatment actions and allocating effort is risk control, not risk identification.
Reporting the residual risk level of each product area to stakeholders in the weekly test progress report.
Incorrect. Reporting residual risk is part of risk monitoring and test reporting, performed once risks are known and testing is under way.
Risk identification techniques include expert interviews, independent assessments, retrospectives and lessons learned from earlier projects, risk workshops (for example FMEA), brainstorming and checklists. Rating likelihood and impact, choosing mitigation actions and reporting residual risk belong to risk assessment, risk control and risk monitoring respectively, which are later steps in risk management.
The test manager of a video streaming service assesses product risks using a risk matrix in which likelihood and impact are each rated as low, medium or high. What does this qualitative approach provide, and what is its main limitation?
It places the risks into relative priority bands so that test effort can be allocated to the most important ones, but the ratings are subjective judgements and cannot meaningfully be combined arithmetically or compared across projects.
Correct. The matrix gives relative prioritization on an ordinal scale; its subjective, non-numeric nature is the limitation of qualitative assessment.
It produces a numeric risk score for each item that can be summed into a single overall project risk figure, but it depends on historical defect and cost data that most projects do not have available in the quality needed for the figure to be trusted.
Incorrect. This describes quantitative risk assessment; a low/medium/high matrix does not yield genuine numeric scores that can be summed.
It records the agreed owner and the planned mitigation action for every risk item, but it cannot show the relative priority of one risk against another until test execution has started and defect data is available.
Incorrect. Showing relative priority before execution is precisely what the matrix does; recording owners and actions belongs to risk control documentation.
It classifies each item as either a project risk or a product risk, but it gives no indication of impact, so the impact of each item has to be established separately in a quantitative cost-based model before test effort is allocated.
Incorrect. The matrix is built from likelihood and impact ratings; it does not exist to separate project risks from product risks.
A risk matrix supports qualitative risk assessment: it places risks into relative priority bands so that test effort can be allocated to the most important ones. Because the ratings are subjective expert judgements on an ordinal scale, the resulting risk levels cannot meaningfully be treated as numbers, aggregated arithmetically or compared across projects.
During the first two sprints on an HR payroll platform the team ran detailed…
During the first two sprints on an HR payroll platform the team ran detailed risk workshops and produced a thorough risk register. By sprint six the register has not been touched, although several new payroll rules have been added to the product. The test manager recognises a well-known difficulty of risk-based testing. Which difficulty is this, and what is the recommended countermeasure?
Keen beginnings: risk analysis is performed enthusiastically at the start and then neglected. Risk analysis should be treated as a continuous activity, with the register revisited and updated at regular planned points throughout the project.
Correct. Initial enthusiasm that fades is the keen beginnings difficulty, and the recommended remedy is to make risk analysis continuous rather than a one-off event.
Key risks are being missed: too narrow a group contributed to the analysis. A wider range of stakeholders should be involved and more than one identification technique should be applied in each analysis session so that blind spots are covered.
Incorrect. This is a genuine risk-based testing difficulty, but it concerns the breadth of the original analysis, not the failure to maintain the register over time.
The analysis was carried out by people without sufficient payroll domain knowledge. Risk identification should be handed over to the business stakeholders who own the payroll rules and repeated at the start of each major release.
Incorrect. Nothing in the situation suggests the wrong people performed the analysis; the register was well produced but never maintained.
This is a test monitoring shortcoming rather than a risk analysis difficulty. The risk register should be replaced by product risk coverage metrics that are published in each sprint test progress report and reviewed by the team.
Incorrect. Reporting coverage metrics does not substitute for maintaining the risk analysis, and abandoning risk analysis after a strong start is a recognised risk-based testing difficulty in its own right.
This is the difficulty known as keen beginnings: risk analysis is done enthusiastically at the start of a project and then abandoned as the project proceeds. The countermeasure is to treat risk analysis as a continuous activity, revisiting and updating the risk register at regular planned points so that it keeps pace with changes to the product.
The test manager of an online travel insurance platform is explaining product…
The test manager of an online travel insurance platform is explaining product risk treatment to the project board. The board wants to understand, in plain terms, how testing relates to the other risk treatment options available and on what basis one option is chosen rather than another. Which statement correctly describes these options and the criteria for choosing between them?
Testing is the primary risk-mitigation measure and is selected where the risk can be reduced by finding defects before release; a contingency plan is prepared where the risk cannot be reduced but its consequences can be limited; transfer passes responsibility to another party; and acceptance is chosen where the cost of treatment outweighs the exposure.
Correct. This is the syllabus distinction: mitigation through testing reduces the risk, a contingency plan limits its consequences, transfer moves responsibility elsewhere, and acceptance tolerates the exposure when treatment is not worth its cost.
Testing is the primary risk-mitigation measure and is selected for each identified risk without further analysis; a contingency plan is the schedule buffer added to the test plan so that slipped test cycles can still finish; transfer means moving the affected feature to a later release; and acceptance is the sign-off the board gives to the completed risk analysis.
Incorrect. Testing is chosen where a risk can actually be reduced by it, a contingency plan addresses the consequences of a risk rather than schedule slack, transfer concerns responsibility rather than release scope, and acceptance is a treatment decision, not an approval of the analysis.
A contingency plan is the primary risk-mitigation measure and is prepared for each high-priority risk; testing is selected only where a risk is rated low and the cost of a contingency plan would not be justified; transfer means reassigning the risk to a different test level in the same project; and acceptance is the decision to stop testing once exit criteria are met.
Incorrect. Testing, not contingency planning, is the primary mitigation measure, and it is directed at the higher risks; transfer involves another party rather than another test level, and meeting exit criteria is a completion decision rather than risk acceptance.
Transfer is the primary risk-mitigation measure and is applied by delegating the affected functionality to the supplier that built it; testing is the contingency plan invoked where transfer is not possible; a contingency plan covers risks the supplier declines; and acceptance is the state a risk reaches once its tests have passed.
Incorrect. Transfer is one option among several rather than the primary measure, testing is a mitigation measure rather than a contingency plan, and a risk whose tests have passed has been mitigated rather than accepted.
Testing is the primary risk-mitigation measure available to a test manager: it reduces a product risk by finding defects before release. Where a risk cannot usefully be reduced in that way, three further treatment options exist. A contingency plan limits the consequences should the risk materialize, risk transfer passes responsibility for the risk to another party such as a supplier or an insurer, and risk acceptance tolerates the exposure without further treatment. The criteria for choosing are whether the risk can be reduced by testing, whether its consequences can be limited if it occurs, whether another party is better placed to carry it, and how the cost of treatment compares with the exposure.
A test manager at an agricultural IoT company is setting up an analytical-based test process improvement approach and wants the team to use the terms measure, metric and indicator precisely. Which statement is correct?
A measure is the value assigned to an attribute, a metric is the measurement scale together with the method of taking the measurement, and an indicator is a measure used to evaluate an attribute against a target; good indicators are effective, efficient and predictable.
Correct. This matches the definitions and the three qualities expected of a good indicator.
A measure is the scale and the method used to take a measurement, a metric is the single value that results from applying that scale, and an indicator is the target value the result is compared against; good indicators are objective, repeatable and independently auditable.
Incorrect. Measure and metric are transposed here, an indicator is not the target itself, and the three qualities listed are not the ones given for indicators.
A measure is a raw value collected from the test process, a metric is the trend that value follows over successive iterations, and an indicator is the control action triggered whenever the trend deviates; good indicators are timely, visible and inexpensive to collect.
Incorrect. A metric is a scale and method rather than a trend, and an indicator is a means of evaluation rather than a control action.
Measure and metric are interchangeable terms for any value that has been collected during testing, while an indicator is any chart or dashboard included in the test progress report; good indicators are effective, efficient and predictable.
Incorrect. Although the three indicator qualities are stated correctly, measure and metric are distinct terms and an indicator is not simply a chart.
A measure is the value assigned to an attribute by a measurement. A metric is the measurement scale together with the method used to take the measurement. An indicator is a measure or combination of measures used to evaluate or estimate an attribute against a defined target. Good indicators are effective, efficient and predictable.
System testing of a smart-parking release has finished. As part of the test…
System testing of a smart-parking release has finished. As part of the test completion activities the test environment is being handed back, the testware and test data are being archived, and the test manager is now writing the test completion report for that test level. What does this work product contain, and who is it addressed to?
It summarises the testing performed for the completed test level or project, including coverage, results, residual risks and lessons learned, and is addressed to stakeholders so that they can judge the quality of the test object.
Correct. The test completion report is a retrospective summary produced at the end of a test level, iteration or project for the benefit of stakeholders.
It is the archive index recording where the testware, the test data and the test environment configurations have been stored, so that the maintenance team can retrieve and reuse them during the next release of the smart-parking platform.
Incorrect. Archiving the testware and recording its location is a separate test completion activity; the report summarises the testing rather than cataloguing the stored assets.
It documents the scope, the approach, the resources and the schedule of the testing that is still to be performed, so that stakeholders can approve the planned effort and its budget before test execution begins.
Incorrect. Scope, approach, resources and schedule are the content of a test plan, produced before testing rather than at test completion.
It lists every executed test case with its individual pass or fail outcome and the defect reports raised against it, so that the developers can reproduce and correct each remaining failure before the test level is closed.
Incorrect. Detailed per-case outcomes belong in the test execution log and defect reports, which serve developers rather than summarising the level for stakeholders.
The test completion report summarises the testing performed for a completed test level, iteration or project. It describes what was covered, the results obtained, defects and residual risks, and lessons learned, and it is addressed to stakeholders so that they can judge the quality of the test object and the effectiveness of the testing. It is produced alongside the other test completion activities, such as archiving the testware and handing back the test environment, but it is not itself the archive record.
A supplier of airport baggage-handling systems wants to introduce a new test management tool across its five engineering sites. Which approach reflects good practice for tool introduction?
Assess the organization's maturity and needs, run a proof of concept on the shortlisted candidates, then a pilot project to evaluate the tool in real conditions, and only then roll it out gradually across the sites.
Correct. Needs assessment, proof of concept, pilot and then incremental rollout is the recommended order for tool introduction.
Assess the organization's maturity and needs, purchase enterprise licences for all five sites at once, then run a pilot project in parallel with the rollout so that its lessons can be applied while adoption is under way.
Incorrect. Committing to enterprise licences and rollout before the pilot removes the chance to act on what the pilot reveals.
Shortlist the candidate tools on functionality and licence cost alone, deploy the chosen tool to all five sites simultaneously so that everyone learns it together, and run a proof of concept afterwards to confirm the choice.
Incorrect. A proof of concept run after deployment cannot influence the selection, and simultaneous rollout to all sites forgoes the pilot stage.
Run a proof of concept on the tool the sites already use, migrate the existing testware into the new tool straight away, and then let each site decide independently whether to adopt it once migration is finished.
Incorrect. The proof of concept must target the candidate tools, and migrating testware before adoption is agreed inverts the sequence.
Good practice for introducing a test tool is to start from an assessment of the organization's maturity and needs, run a proof of concept against the shortlisted candidates to confirm they work in the intended technical context, then run a pilot project to evaluate the tool in real conditions and gather lessons, and only afterwards roll the tool out incrementally across the organization.
A test manager at a freight shipping company is preparing the business case for a new test automation tool. Which TWO of the following are factors that influence the tool decision according to the syllabus? (Select TWO)
Applicable regulations and security constraints, such as where consignment test data and test results may legally be stored.
Correct. Regulations and security are one of the four listed factors influencing the tool decision.
How well the tool fits the existing software landscape, including the current CI pipeline and defect management system.
Correct. Fit with the existing software landscape is one of the four listed factors influencing the tool decision.
The defect detection percentage that the test team achieved on the two most recent releases of the freight shipping platform.
Incorrect. Defect detection percentage measures past test effectiveness; it is not one of the listed tool decision factors.
The number of test cases that are currently maintained in spreadsheets rather than in any dedicated test management tool.
Incorrect. This describes the current state of the testware inventory rather than a factor the syllabus lists for the tool decision.
The test levels and the test types at which the previously purchased tool happened to be used by the shipping teams.
Incorrect. Historical usage of an earlier tool is background information, not one of the listed decision factors.
The syllabus identifies four factors influencing the decision on which test tool to use: applicable regulations and security constraints, financial aspects, stakeholder requirements, and the existing software landscape into which the tool must fit. Measures of the current test effort, such as defect detection percentage or the number of spreadsheet-based test cases, describe the present situation but are not the decision factors listed.
A test manager taking over an electric-vehicle charging network programme begins by identifying the test stakeholders. Which statement best describes who counts as a test stakeholder and why identifying them early matters?
Anyone with an interest in the testing or affected by its results, such as developers, product owners, operations staff, customers and regulators; identifying them early lets the test manager establish their information needs and expectations and plan their involvement.
Correct. Test stakeholders are defined broadly, and early identification allows their needs and involvement to be reflected in the test approach and reporting.
The parties who fund the test effort and hold the budget for it, such as the programme sponsor and the finance function; identifying them early lets the test manager secure the funding and the specialist staffing that are needed before the master test plan can be baselined.
Incorrect. Sponsors are stakeholders, but restricting the definition to those who pay excludes operations, customers, regulators and others affected by the results.
The members of the test team together with their line management and the test architects supporting them; identifying them early lets the test manager assign roles and responsibilities inside the team before test analysis and test design work begins.
Incorrect. This limits stakeholders to the test organization itself, whereas most stakeholders sit outside the test team.
The parties who formally approve the entry and exit criteria for each test level, such as the product owner and the release manager; identifying them early lets the test manager schedule the approval gates in the test plan and the test schedule.
Incorrect. Approvers are one group of stakeholders, but many stakeholders have information needs without holding any approval authority.
Test stakeholders are all those who have an interest in the testing or who are affected by its results, including developers, product owners, operations and support staff, customers and regulators. Identifying them early allows the test manager to establish their information needs and expectations and to define their involvement, so that the test approach and reporting can be shaped accordingly.
A test manager at a maritime port automation supplier has an approved…
A test manager at a maritime port automation supplier has an approved organizational test strategy and is now preparing the testing for a specific terminal modernisation project. A colleague asks what the project's test approach is and what it should comprise. Which statement best describes a test approach?
It is the implementation of the test strategy for a specific project or release, defining the test techniques and test levels to be applied, the entry and exit criteria to be met, and the degree of test automation to be achieved.
Correct. The syllabus defines the test approach as the implementation of the test strategy for a particular project, comprising decisions on techniques, levels, entry and exit criteria, and degree of automation.
It is the long-term, organization-wide document describing the generic testing principles, the test process and the standards applied across all of the projects, and it is revised only when the company's quality policy itself changes.
Incorrect. That describes the organizational test strategy. The test approach is what implements that strategy in the context of one specific project.
It is the part of the test plan that lists the test tasks and their dependencies, the effort estimate for each task, and the assignment of named testers to each of the test levels planned for the current terminal release.
Incorrect. Scheduling, estimation and staffing are test planning outputs. They are informed by the test approach but are not what the test approach comprises.
It is the set of measurements collected while testing is running, such as the coverage achieved, the defect density and the residual risk, which are used to report test progress against the criteria agreed in advance.
Incorrect. Those measurements belong to test monitoring and test reporting. They report against criteria that the test approach defined in advance.
The test approach is the implementation of the test strategy for a particular project or release. It tailors the generic strategy to the project context and typically defines the test techniques to be used, the test levels and test types, the entry and exit criteria, and the intended degree of test automation. The long-lived organization-wide document is the organizational test strategy; staffing and scheduling belong to test planning; measurements gathered during execution belong to test monitoring.
A test manager on a clinical trial data capture system is reviewing the exit…
A test manager on a clinical trial data capture system is reviewing the exit criteria proposed by the team for the system test level, which is planned to end on 14 November. The proposed criteria are: EC1: All high-priority test cases have been executed. EC2: The system is stable enough for the trial sponsor. EC3: No more than five severity-2 defects remain open, taken from the defect management tool on 14 November, with no severity-1 defects open. EC4: All 4,000 planned regression test cases pass before the end of the project. The test manager applies the S.M.A.R.T. attributes to each criterion before taking the set to the sponsor. Which analysis correctly identifies the weakest criterion, the attributes it fails and an appropriate correction?
EC2 is neither Specific nor Measurable, because 'stable enough' names no property and no threshold that can be evaluated; it should be restated with a measurable indicator and value, for example that the mean time between failures observed in the last full regression run on 14 November is at least 40 hours.
Correct. EC2 is the only criterion with no measurable property at all, so agreement on whether it is met would be a matter of opinion. Naming an indicator, a threshold and the date makes it specific, measurable and time-bound.
EC1 fails the Measurable attribute, because it does not state what percentage of the high-priority test cases must be run; it should be restated to require that at least 95 per cent of the high-priority test cases have been executed and their results recorded in the tool by the planned exit date of 14 November.
Incorrect. 'All' is an unambiguous 100 per cent, so EC1 is measurable. Its only real gap is that it is not time-bound, and lowering the target to 95 per cent weakens the criterion rather than making it S.M.A.R.T.
EC3 fails the Time-bound attribute, because it fixes a fixed calendar date instead of a project milestone; it should be restated so that the severity-1 and severity-2 defect counts are taken from the defect management tool at the end of the system test level, whenever that level actually completes.
Incorrect. A fixed date is exactly what makes a criterion time-bound. Replacing it with 'whenever the level completes' would remove the time boundary and make the criterion self-fulfilling.
EC4 fails the Specific attribute, because it does not name the test level to which the regression suite applies; it should be restated to require that all 4,000 planned regression test cases pass at every test level, from component testing through to acceptance testing, before the release.
Incorrect. EC4's real weaknesses are achievability, since a 100 per cent pass rate on 4,000 cases is rarely attainable, and its vague deadline. Extending it to every test level makes it even less achievable.
S.M.A.R.T. exit criteria are Specific, Measurable, Achievable, Relevant and Time-bound. EC2 is the weakest: 'stable enough for the trial sponsor' names no measurable property and no threshold, so it is neither specific nor measurable and cannot be objectively evaluated on the exit date. Restating it with a defined measure and threshold, tied to the same exit date, makes it assessable. EC3 is already well formed. EC1 is specific and measurable ('all' means 100%) but lacks a time frame. EC4's weakness is achievability and its vague deadline, not specificity.
A satellite ground station software programme is delivered by three teams in…
A satellite ground station software programme is delivered by three teams in three countries. The test manager is facilitating the retrospective for the latest increment using the five-step retrospective process. The two previous retrospectives went badly: participants at the two smaller sites stayed silent, and the discussion turned into an argument over which site had introduced the defects that escaped to the acceptance test level. The test manager has set the stage, has agreed with all three sites that the session is blame-free, and is now moving into the 'Collect data' step. Which action best implements that step in this situation?
Gather objective measurements for the increment before the session, such as defect detection percentage per test level, planned against actual test execution, and test environment downtime, and combine them with input collected anonymously from every site, then present the result as facts about the process rather than about any team.
Correct. This is the 'Collect data' step: it builds a shared factual basis, and gathering input anonymously and framing it as process data lets the quieter sites contribute without the session sliding back into blame.
Open the session by restating the working agreement with all three sites, confirming once more that the session is blame-free, that everything said within it stays inside the group, and that the outcome is owned collectively by the programme, so that the participants at the two smaller sites feel safe enough to speak openly about the increment.
Incorrect. Establishing safety and working agreements is the 'Set the stage' step, which the test manager has already completed. It does not produce the factual picture that the next step requires.
Work through the observations with the group in order to establish the root causes of the two largest deviations seen in the increment, so that the discussion moves from the symptoms that the teams noticed during test execution to the underlying weaknesses in the test process itself, and record the causes agreed.
Incorrect. Root cause analysis is the 'Generate insights' step. It can only be done well once the data has been collected, and doing it first invites the speculation that fuelled the earlier arguments.
Ask each of the three sites to propose the two process changes that it believes would most improve the next increment, then have the whole group vote on the combined list of proposals so that the two changes with the highest expected benefit are carried forward as improvement actions for the increment.
Incorrect. Proposing and selecting changes is the 'Decide on improvement actions' step. Jumping to actions without collected data produces improvements based on impressions rather than evidence.
In the 'Collect data' step the group builds a shared, factual picture of what actually happened in the increment before any interpretation begins. Where earlier sessions turned into blame, the data must be objective, attributed to the process rather than to individuals, and gathered so that quieter participants can contribute without exposure. Agreeing ground rules belongs to 'Set the stage', root-cause discussion belongs to 'Generate insights', and selecting changes belongs to 'Decide on improvement actions'.
A test manager on a water utility SCADA replacement has 100 person-days of test…
A test manager on a water utility SCADA replacement has 100 person-days of test effort available for the release. The team has performed risk analysis on the five features in scope, rating likelihood and impact on a scale of 1 (lowest) to 5 (highest). The agreed method is to calculate each feature's risk level as likelihood multiplied by impact, and then to allocate the available effort to the features in direct proportion to their risk levels. Pump control: likelihood 4, impact 5 Alarm escalation: likelihood 3, impact 5 Historian archiving: likelihood 2, impact 3 Operator dashboard: likelihood 3, impact 2 Report export: likelihood 1, impact 3 Using the agreed method, how much effort should be allocated to Alarm escalation and to Historian archiving?
Alarm escalation 30 person-days and Historian archiving 12 person-days, because the five risk levels multiply out to a total of 50 and each single risk point is therefore worth two person-days of the 100 available.
Correct. Products are 20, 15, 6, 6 and 3, a total of 50. At 2 person-days per risk point this gives 40, 30, 12, 12 and 6 person-days, so Alarm escalation gets 30 and Historian archiving gets 12.
Alarm escalation 26 person-days and Historian archiving 16 person-days, because the five risk levels add up to a total of 31 and the whole of the available effort is shared out in proportion to each feature's combined rating.
Incorrect. This adds likelihood and impact instead of multiplying them, giving 9, 8, 5, 5 and 4. The agreed method is likelihood multiplied by impact, which spreads the effort far more sharply towards the highest-risk features.
Alarm escalation 28 person-days and Historian archiving 17 person-days, because the five impact ratings come to a total of 18 and the effort should follow the potential damage rather than the probability.
Incorrect. This allocates on impact alone and ignores likelihood, which contradicts the agreed method. Risk level in the syllabus combines both the likelihood of the failure and its impact.
Alarm escalation 20 person-days and Historian archiving 20 person-days, because the five features each take an equal baseline share of 20 and the risk level is used only to sequence the execution order.
Incorrect. An equal split is not proportional allocation. Risk level in a risk-based approach determines the depth and extent of testing for each feature, not only the order in which it is executed.
Risk levels are 20, 15, 6, 6 and 3, summing to 50. Each risk point is therefore worth 100 / 50 = 2 person-days. Alarm escalation receives 15 x 2 = 30 person-days and Historian archiving receives 6 x 2 = 12 person-days. The remaining features receive 40, 12 and 6 person-days respectively.
A national identity card issuance system is in week 3 of a 6-week system test…
A national identity card issuance system is in week 3 of a 6-week system test level. The test manager reviews the monitoring data for the week: - Planned cumulative test execution 600 test cases, actual 310. - The test environment was unavailable for 22 of the 60 scheduled test hours, in every case because the nightly deployment failed and had to be repaired the following morning. - 41 per cent of the failed test results raised this week were reclassified after investigation as environment faults rather than product defects. - The biometric enrolment regression pack had to be re-run twice after mid-run environment restarts. Which control directive should the test manager issue first?
Require that every nightly build pass a deployment verification test before it is handed to the test team, assign dedicated environment support for the remaining weeks, then re-plan the execution schedule against the recovered capacity and escalate to the project manager if stability is not restored within one week.
Correct. The environment instability is the single cause behind the lost hours, the false failures and the repeated regression runs. Gating the build and resourcing the environment removes that cause, and re-planning plus escalation handles the shortfall that has already accumulated.
Add two further testers to the execution team for the remaining three weeks and extend the daily test window, so that the shortfall of 290 test cases is recovered by raising the execution rate well above the original plan, and report the revised burn-up each week until the plan is back on track and the trend is visible to the sponsor.
Incorrect. The shortfall was caused by the environment being unavailable, not by a lack of testers. More testers would contend for the same unstable environment and would produce more results needing reclassification.
Increase test progress reporting from weekly to daily and add environment downtime, result reclassification rate and re-run count to the report, so that stakeholders can see the trend well before the planned end date of the test level and can decide for themselves what corrective action is warranted.
Incorrect. Better reporting is test monitoring, not test control. It improves visibility of the problem but issues no corrective action, so the environment continues to consume test hours.
Relax the exit criteria for the system test level, reducing the required execution coverage in line with the rate actually achieved so far and limiting the biometric enrolment regression pack to a single run, so that the planned end date of the test level can still be met without changing the way of working.
Incorrect. Lowering the exit criteria accepts the shortfall instead of correcting its cause, silently increases residual risk on an identity issuance system, and in any case cannot be decided by the test manager alone.
Test control means taking corrective actions when monitoring shows a deviation. Here the data all point to one cause: an unstable test environment produced by unverified nightly deployments. The directive must address that cause, so that the wasted execution and the false failures stop, and the execution schedule must then be re-planned against what remains achievable, with escalation if stability is not restored. Adding testers, increasing reporting frequency or relaxing exit criteria all leave the cause untouched.
A container terminal operator is replacing its terminal operating system. The…
A container terminal operator is replacing its terminal operating system. The delivered service is a system of systems: the new terminal operating system itself, which the operator's supplier builds; the crane control units, whose firmware is supplied by a plant manufacturer under a contract that gives the operator no right to test the firmware or access its source; the national customs authority gateway, which offers only a certified conformance sandbox with limited availability; and a truck appointment service consumed as software as a service, whose provider publishes an interface specification and monthly test reports. The test manager must define the test scope for the programme and present it for approval. Which scope definition and justification is the most appropriate?
In scope: the terminal operating system and every interface to the third-party components, exercised through the customs conformance sandbox and through service virtualization where the real component is unavailable. Out of scope: the internal behaviour of the crane firmware, the customs gateway and the appointment service, for which supplier test reports and certificates are reviewed as evidence and the remaining risk is recorded in the test plan and reported.
Correct. The interfaces carry the integration risk the operator owns and can test, while the internal behaviour of components it has no right to test is placed out of scope explicitly, with supplier evidence and a documented residual risk rather than silent omission.
In scope: only the components the operator's own supplier builds, because the organization has no contractual right to test any third-party software. The interfaces are covered by each supplier's own acceptance testing and by the monthly test reports published by the appointment service provider, so they can be treated as already verified and left out of scope, with the certified customs sandbox used only if spare capacity happens to be available late in the programme.
Incorrect. A supplier can only test its own side of an interface against its own assumptions. Leaving the interfaces untested puts the highest-risk area of a system of systems outside the scope entirely.
In scope: everything the end-to-end business process touches, including the internal logic of the crane firmware and of the customs gateway, because the terminal operator is accountable for the whole service to its customers and cannot delegate that accountability to its suppliers. The contractual restrictions are recorded as an obstacle for the procurement function to resolve, while test design proceeds on the assumption that the access will be granted in time.
Incorrect. Accountability for the service does not create the access, source code or contractual right needed to test supplier internals. A scope that cannot be executed is not a scope definition, and it hides the real dependency instead of managing it.
In scope before release: only the components the team owns, with the end-to-end flows across all four systems moved into post-release monitoring in production, because integration behaviour between independently operated systems can realistically only be observed once every supplier is connected in the live environment. A staged rollback plan and enhanced production alerting are documented in the test plan as the compensating measures for the deferred checks.
Incorrect. Deferring all integration verification to production accepts integration failures on a live terminal as the detection mechanism. The conformance sandbox and service virtualization make meaningful pre-release interface testing feasible.
Scope definition for a system of systems must separate what the organization can and must test from what it cannot test but still depends on. The terminal operating system and every interface to the third-party components are within the operator's control and carry the integration risk, so they are in scope, using the conformance sandbox and service virtualization where the real component is not available. The internal behaviour of components the operator cannot test is out of scope, but the dependency is not ignored: supplier evidence is reviewed and accepted, and the residual risk is recorded in the test plan and reported to stakeholders.
A health insurer runs two programmes that share the single integrated test…
A health insurer runs two programmes that share the single integrated test environment connecting to the claims mainframe. Programme A replaces the prescription reimbursement rules engine; a change in national reimbursement law makes 1 October an immovable date, and the risk analysis rates incorrect reimbursement calculations as the highest product risk in the portfolio. Programme B redesigns the pharmacy portal user interface; it is commercially valuable, has no external deadline, and its highest-rated risks are usability and browser compatibility. Both programmes are funded from the same budget line and both have asked for continuous access from July onwards. The test manager owns the allocation decision. Which allocation and analysis of its consequences is the most appropriate?
Give Programme A priority access in blocks aligned to its 1 October milestones, and give Programme B fixed scheduled windows plus a virtualized instance for the usability and compatibility testing that does not need the mainframe. Record in both test plans that B's end-to-end testing is compressed, state the resulting residual risk, and escalate the environment capacity shortage as a project risk.
Correct. The allocation follows the risk levels and the immovable regulatory date, uses virtualization to cover the work that does not need the shared connection, and makes the consequence for Programme B visible and owned rather than absorbed silently.
Split the environment evenly by alternating weeks between the two programmes, because they are funded from the same budget line and an even split is the allocation that can be defended to both sponsors without reopening the risk analysis. Any resulting schedule impact is then borne symmetrically, is visible in each programme's own plan, and needs no separate escalation to the steering committee at this stage.
Incorrect. Equal funding is not equal risk. An even split gives the same capacity to a usability redesign with no deadline as to a highest-risk change with a statutory date, and the symmetric impact falls hardest where it can least be absorbed.
Give Programme B the environment first because its release is smaller and can be cleared in a few weeks, which then frees continuous uninterrupted access for Programme A from August onwards. Completing the shorter piece of work first minimises the total waiting time across the two programmes and removes the contention before the regulatory milestone approaches in the autumn.
Incorrect. Minimising aggregate waiting time optimises the wrong objective. It spends Programme A's contingency early and leaves no recovery time if the reimbursement rules engine testing uncovers serious defects before 1 October.
Operate the environment on a first-come, first-served booking basis run by the environment team, because prioritising between two independently sponsored programmes is a governance decision rather than a test management one. Contention that cannot be resolved between the two test leads is then escalated to the steering committee when it arises, with the booking log as evidence.
Incorrect. Booking order is unrelated to risk, and escalating only after contention occurs means the conflict surfaces once slots are already lost. The test manager holds the risk information needed to propose the allocation.
Resource allocation between contending programmes is a risk-based decision, not a fairness decision. Programme A combines the highest product risk with a deadline that cannot move, so it takes priority for the shared environment. Programme B's dominant risks, usability and browser compatibility, do not require the mainframe connection for most of their testing, so scheduled windows plus a stubbed or virtualized instance cover much of its need. The consequences must be made explicit: B's end-to-end testing is compressed, and that residual risk is documented in both test plans and the capacity constraint escalated as a project risk.
A mining equipment manufacturer is developing the collision avoidance function…
A mining equipment manufacturer is developing the collision avoidance function for its autonomous haul trucks. The hazard and risk analysis performed under IEC 61508 has assigned safety integrity level SIL 3 to the function that stops a truck when an obstacle is detected in its path. The same embedded product also contains a fleet telemetry and shift reporting function, which the analysis assigns no safety integrity level. The test manager must now analyse what the assigned SIL means for the test plan before the plan is submitted for assessment. Which analysis is correct?
SIL 3 drives both the techniques and the rigour: the plan must specify strong structural coverage of the safety function together with techniques such as boundary value analysis, equivalence partitioning, fault injection and static analysis, a defined independence for the verification activities, and full traceability from each safety requirement to its test cases and retained results. The telemetry function may be tested with lower rigour provided the segregation is justified.
Correct. The standard's recommended techniques and the required independence strengthen as the SIL rises, and the evidence chain from safety requirement to retained result is what the assessor examines. Applying lower rigour to a non-safety function is permitted only with a documented segregation argument.
SIL 3 governs the documentation and record keeping rather than the technical work: the choice of test techniques stays at the test manager's discretion exactly as before, and the plan changes only by adding a formal test summary report for each test level, a signed traceability matrix covering the safety requirements, and long-term archiving of all test results and defect records so that the assessor can reconstruct what was done, by whom and in what order at any point in the project.
Incorrect. The standard recommends specific techniques and measures at each integrity level, so the choice of techniques is constrained, not discretionary. Documentation records the rigour applied; it does not substitute for it.
SIL 3 applies uniformly to the whole product, requiring full statement coverage of every line of code including the telemetry and shift reporting function, so the plan should size the test effort from the size of the code base rather than from the risk analysis, since any function shipped inside a safety-related product is itself safety-related and must carry the same evidence and the same degree of verification independence when presented to the assessor.
Incorrect. The integrity level is assigned to a safety function, not to an entire code base, and effort follows the risk analysis. Spreading uniform rigour across non-safety code consumes effort that the safety function needs.
SIL 3 requires the safety function to be verified solely by an accredited external certification body, so the plan should exclude that function from the project's own component and integration test levels and reference the certification body's schedule, scope and techniques as the coverage for it instead, with the internal team concentrating its remaining effort on the telemetry function and on system integration once the certificate has been issued.
Incorrect. The standard requires a defined degree of independence for verification, which can be met internally, and external assessment reviews the supplier's own evidence. Removing the safety function from the project's test levels would leave the assessor with nothing to assess.
IEC 61508 is a risk-based safety standard whose recommendations for techniques and measures become stronger as the safety integrity level rises. At SIL 3 the standard drives rigorous structural coverage of the safety function, supporting techniques such as boundary value analysis, equivalence partitioning, fault injection and static analysis, a defined degree of independence for the verification and validation activities, and complete bidirectional traceability from safety requirements to test cases and retained results for the assessor. Non-safety functions in the same product may be tested with lower rigour, but the segregation argument has to be justified and documented.
A warehouse robotics programme depends on a motion controller board available…
A warehouse robotics programme depends on a motion controller board available from a single specialist supplier. Two facts have just been added to the risk register. First, the board now has a confirmed lead time of 26 weeks, against the 8 weeks assumed when the release plan was drawn up. Second, a credit report shows the supplier is in financial distress and may not survive the year. The programme manager asks the test manager whether the test team can mitigate this risk by testing it harder, pointing out that the risk register currently has no mitigation entry at all. Which analysis and treatment should the test manager present?
It is a project risk that testing cannot influence, because neither its likelihood nor its impact arises from a defect in the product. The treatment is contingency planning, qualifying an alternative controller and keeping the control software portable, combined with transfer through contract terms and acceptance of the residual delay. The test plan records the dependency and its effect on test environment availability.
Correct. Test activities can only reduce product risk arising from defects. A supply chain risk is treated by contingency, transfer and acceptance, and the test manager's contribution is to make the consequences for the test environment and schedule visible.
It is a product risk, because switching to a substitute controller would change the system's timing and interrupt behaviour. The correct mitigation is therefore to extend hardware compatibility and timing testing across every candidate board now, which reduces the likelihood that the supplier situation affects the quality of the release and gives the programme a tested fallback some weeks before the current lead time expires.
Incorrect. This confuses a consequence of the risk with the risk itself. Compatibility testing is worth doing as part of the contingency plan, but it does not change the supplier's finances or the 26-week lead time, so it does not mitigate this risk.
It is a project risk, and the correct treatment is mitigation by testing earlier: pulling the integration test level forward and running it against simulated controller inputs reduces the impact of the delay, because the defects will already have been found by the time the boards arrive and only a short confirmation run on the real hardware will remain before the release can be approved by the programme board.
Incorrect. Earlier integration testing against simulators is useful preparation, but it neither reduces the likelihood of the shortage nor removes the dependency on physical boards for final verification, so the delay impact is unchanged.
It is a project risk that should be accepted and then closed in the register, because a risk that the test organization cannot influence through any test activity falls outside test management and belongs to the procurement function to handle under its own supplier governance process, with the test team simply informed of the revised delivery date once procurement has confirmed a new supply date.
Incorrect. Acceptance is a legitimate treatment, but an accepted risk stays in the register and stays monitored. Closing it removes the trigger for the contingency plan and hides a dependency that directly affects test environment availability.
This is a project risk in the supply chain. Its likelihood is driven by the supplier's finances and manufacturing capacity and its impact by delivery dates, and neither is influenced by any test activity, because no defect in the software causes it. Testing therefore cannot mitigate it. The appropriate treatments are contingency planning, such as qualifying a second board and keeping the software portable across controllers, combined with transfer through contractual or insurance arrangements and acceptance of the residual delay. The risk stays owned and monitored, and the test plan records the dependency and its effect on test environment availability.
An online betting exchange has an approved organizational test strategy written…
An online betting exchange has an approved organizational test strategy written three years ago. It assumes a sequential lifecycle with a formal system test phase completed before each release, and it requires the shared test environment to be loaded with a full anonymised copy of production data before that phase begins. The new exchange platform breaks both assumptions: it is delivered by continuous delivery with several production releases a day, and the gambling regulator's licence conditions now forbid copying customer betting histories into any non-production environment, even anonymised. The test manager must analyse the strategy against the project and propose how to proceed. Which course of action is the most appropriate?
Keep the organizational strategy as the reference and record two justified deviations: the phase-based system test level is replaced by a continuously executed risk-based automated regression suite meeting the same coverage objectives, and the production data copy is replaced by generated synthetic data plus a small masked subset. Each deviation is documented in the test plan with its rationale, residual risk and approval by the strategy owner, and both are fed back so the strategy can be updated.
Correct. Each deviation is justified against the objective the original rule served, approved by the strategy owner and recorded with its residual risk, and the feedback loop keeps the organizational strategy relevant for the projects that follow.
Grant the programme a full exemption from the organizational strategy and let it define its own standalone test strategy for its duration, because a project that cannot comply with two central rules cannot meaningfully claim to follow the strategy at all. The exemption is approved once by the quality board and recorded in the programme charter, and the organizational strategy is then applied unchanged to the next project that starts, so that the exemption stays an isolated case rather than a precedent.
Incorrect. A blanket exemption discards the parts of the strategy the project could still follow and produces no traceable justification per deviation. It also loses the feedback that would keep the organizational strategy from failing the next project the same way.
Adapt the project to the strategy rather than the strategy to the project: batch the continuous releases into a monthly system test phase and apply to the gambling regulator for an exemption allowing a masked production copy to be loaded into the shared environment, because deviating from an approved strategy undermines the consistency and comparability of test results across the organization and weakens the audit position at licence review and at future audits.
Incorrect. Consistency is not an end in itself, and this reshapes the delivery model and asks a regulator to relax a licence condition in order to preserve a document. The strategy exists to serve the projects, not the reverse.
Raise a deviation only for the test data rule, since that constraint is externally imposed by the licence conditions and cannot be argued with, and leave the lifecycle rule untouched, because the delivery cadence is a development concern rather than a testing one; the system test level can stay documented exactly as written and simply be executed once each quarter against the changes accumulated in the shared environment since the previous run.
Incorrect. A quarterly system test level behind several releases a day would let untested changes reach production continuously. The lifecycle model is one of the factors that shapes the test strategy, so it requires its own justified deviation.
The organizational test strategy remains the reference document, and a project does not abandon it because parts of it do not fit. Where the project context genuinely differs, deviations are identified individually, each deviation is justified against the objective the original rule served, and each is documented in the project's test plan with its rationale, its residual risk and formal approval from the strategy owner. Because both differences here are structural rather than one-off, the deviations are also fed back so that the organizational strategy can be updated for future projects.
A pharmaceutical manufacturer is building a new manufacturing execution system…
A pharmaceutical manufacturer is building a new manufacturing execution system for a sterile production line. The system falls under regulated good manufacturing practice, and the electronic batch records it produces must satisfy the regulator's rules on electronic records and signatures; loss of the site's manufacturing licence would halt all production. The board has stated that maintaining regulatory approval takes precedence over the delivery date. The team wants to work in two-week increments, several plant interfaces are still being specified, and realistic batch data is hard to obtain. The test manager is deciding which of the factors that influence the test strategy should dominate. Which analysis is the most appropriate?
The domain and the organizational goals dominate: the regulated domain makes validation evidence, traceability and audit-ready records mandatory and so constrains which test process and techniques are permissible, and retaining the manufacturing licence outranks the delivery date. The remaining five factors still shape the strategy but are chosen within those boundaries, so incremental delivery is acceptable provided each increment produces the required validation evidence.
Correct. Regulatory obligations and the stated organizational goal define what the strategy must achieve, while the lifecycle, resources, interfaces and data situation determine how it is achieved inside that envelope.
Test resources and the SDLC model dominate: the regulatory documentation burden is fixed and cannot be negotiated, so the only variables the test manager can genuinely plan around are the staffing available, the tooling budget and the delivery cadence chosen for the increments. The strategy is therefore built from the team's capacity and its two-week rhythm, with the regulated obligations treated as a constant background condition that the site quality function discharges in parallel.
Incorrect. A fixed obligation is a dominating influence precisely because it constrains everything else. Treating it as background and planning only around the adjustable factors risks producing a strategy that cannot satisfy the regulator.
The project goals and the project type dominate, because a test strategy is tailored to the individual project and to the objectives set for it. The regulated domain is an organizational-level constraint that the company's existing quality management system already satisfies through its standard operating procedures, so it need not shape this project's own test strategy beyond a reference to those procedures in the project test plan and its approval record.
Incorrect. A quality management system defines the framework but does not by itself produce the validation evidence for this system. The domain shapes the test techniques, documentation and traceability required at project level.
The availability of test data and the interfaces with other systems dominate, because validation evidence in this domain is generated from data-driven runs across the connected plant equipment, which makes access to realistic batch data and to the still-unspecified plant interfaces the binding practical constraints on the entire test strategy and on the sequence in which the required evidence can be produced and submitted to the regulator.
Incorrect. Both are genuine constraints on execution, but they determine how the evidence is produced, not what evidence is required. The regulator's requirements and the board's stated goal set that, and they remain in force whatever the data situation.
All seven factors shape a test strategy, but in a heavily regulated context the domain and the organizational goals set the boundaries within which the other five are chosen. The domain makes validation evidence, traceability and audit-ready records mandatory rather than optional, and the organizational goal of retaining the manufacturing licence outranks project schedule goals. Project type, test resources, the lifecycle model, interfaces and test data availability remain real influences, but they are decided inside those boundaries: an incremental lifecycle is acceptable provided every increment still produces the required validation evidence.
A supplier of air traffic flow management software proposes adopting an open…
A supplier of air traffic flow management software proposes adopting an open source test automation framework, extended in house, to replace manual regression testing. The business case submitted to the steering committee reads: licence cost zero, because the framework is open source; one-off setup and framework extension 20 person-days by two senior automation engineers; expected annual saving 400 person-days of manual regression execution; conclusion, the investment repays itself within the first month and delivers a positive return from then on. No other figures appear, and no person is named as responsible for the framework after it goes live. The test manager is asked to evaluate the case. Which evaluation is the most appropriate?
The case is unsound: it counts only non-recurring acquisition cost and omits the recurring costs of script and framework maintenance, dependency upgrades, training and pipeline infrastructure, together with the opportunity cost of two senior engineers withdrawn from risk-based test analysis; the saving also assumes the full manual regression would otherwise always have been run. It should be reworked with annual recurring costs and a named tool owner accountable for the framework's evolution, standards and support.
Correct. Zero licence fee removes one non-recurring cost only. Without recurring costs, opportunity cost and a named owner to keep the framework maintained, the projected return rests on assumptions the case never states.
The case is sound in principle, since an open source licence genuinely removes the acquisition cost from the calculation, but it should be strengthened before submission by benchmarking two further open source frameworks against it on interface coverage and reporting capability, and by extending the evaluation horizon from one year to five years, so that the 20 person-day setup is amortised over a longer period and the committee sees a cumulative saving rather than a first-month payback figure that it is likely to treat as optimistic.
Incorrect. Extending the horizon of a calculation that omits every recurring cost makes the projected return look larger, not more accurate. Benchmarking rival frameworks is a reasonable procurement step, but it does not supply the missing cost categories or the accountable tool owner.
The main flaw is the unit of measurement: expressing the saving in person-days rather than in currency prevents the steering committee from comparing this proposal against the other investments competing for the same capital budget in the same period. Converting the 400 person-days at the loaded engineer rate, adding a 20 per cent contingency to absorb estimation error, and restating the payback in budget months would make the case acceptable in the format the committee already uses for tooling proposals.
Incorrect. Currency conversion changes the presentation of the figures, not their content, and a contingency percentage applied to a benefit figure inflates the benefit rather than supplying the recurring costs and opportunity cost the case leaves out.
The main flaw is that no pilot has been run: the decision should be deferred until a proof of concept on one subsystem has demonstrated that the framework drives the flow management interfaces reliably and produces usable defect reports. Once technical fit has been established in that way, the same cost and saving figures can be carried into the final business case unchanged, because a successful pilot confirms the assumptions on which they rest.
Incorrect. A pilot establishes technical fit, which is a different question from whether the economics hold. Carrying the same figures through unchanged preserves exactly the omissions that make the case unsound, and a pilot says nothing about who will own the framework once it is live.
The case is unsound because it treats acquisition cost as the whole cost. A zero licence fee removes only one non-recurring cost; the recurring costs of maintaining the framework and the test scripts as the product changes, upgrading dependencies, training new testers and running the pipeline infrastructure are missing, as is the opportunity cost of two senior engineers withdrawn from risk-based test analysis. The claimed saving also assumes the full manual regression would otherwise have been executed every time. Finally, no tool owner is named, and without an owner responsible for the framework's evolution, standards, training and support, an in-house or open source tool degrades and the predicted return is never realised.
A national digital land registry is being delivered in eight increments over two…
A national digital land registry is being delivered in eight increments over two years, each increment adding cadastral functions that later increments build on and integrate with. The programme board has asked the test manager to define the set of criteria for the quality gate that every increment must pass before it is accepted into the release baseline. Earlier programmes at this organization used a single criterion, that testing had finished, which allowed defects and untested integrations to accumulate silently across increments until the final release. The test manager must now define criteria that are objectively measurable at the gate date and that protect the increments already accepted. Select TWO criteria that belong in this quality gate set.
No severity-1 defects are open in the functionality delivered by the increment and at most three severity-2 defects are open, each with an agreed and documented workaround, counted from the defect management tool on the gate date.
Correct. The criterion names the measure, the threshold, the source and the date, so it can be evaluated objectively and cannot be argued at the gate, and it forces open issues to be handled explicitly rather than carried forward unseen.
The regression suite covering all previously accepted increments has been executed against the integrated build for this increment with a pass rate of at least 98 per cent, and every failure has been investigated and classified before the gate is assessed.
Correct. In a programme where each increment builds on earlier ones, a regression criterion over the accepted baseline is what stops integration damage accumulating, and requiring every failure to be classified prevents a pass rate being met by leaving failures unexamined.
At least 90 per cent of the test cases planned for this increment have been executed by the gate date, irrespective of their outcome, because execution progress is the most objective and readily available indicator at the point when the gate is assessed.
Incorrect. Execution counted without regard to results measures activity, not product quality. An increment where nine test cases in ten were executed and half of them failed would pass this criterion.
The test effort consumed by the increment is within the estimated number of person-days and the team has recorded no overtime during the increment, demonstrating that the testing was performed in a controlled and sustainable manner.
Incorrect. Effort and overtime describe how the work was run, not whether the product is fit to enter the baseline. An increment can stay within budget precisely because too little testing was done.
Quality gate criteria must be objectively measurable on the gate date and must protect what has already been accepted. A defect criterion expressed in severity thresholds with agreed workarounds, read from the defect management tool on the gate date, gives an unambiguous product quality measure. A regression criterion over all previously accepted increments, executed on the integrated build with a defined pass rate and every failure analysed, protects the accumulated baseline in a multi-increment programme. Execution progress with no reference to results measures activity, not quality, and effort or overtime figures measure the process rather than the product.
Chapter 2 · Managing the Product — 15 questions
A test manager on a museum ticketing programme is explaining to a newly joined…
A test manager on a museum ticketing programme is explaining to a newly joined tester why test reporting takes place while test execution is still running: how often an in-progress test progress report is produced and what it reports the current status against. Which statement describes this correctly?
Test progress reports are produced at regular intervals during the test activity and report the current status of testing against the test plan and the agreed schedule, so that stakeholders can take control actions while there is still time for them to have an effect.
Correct. Progress reporting is periodic by definition, its reference point is the test plan and the test schedule, and its purpose is to keep monitoring and control possible while the activity is still running.
Test progress reports are produced only when an exit criterion has been missed and report the status of testing against the defect database, because reporting at fixed intervals would spend tester effort on information that stakeholders have not asked the test team to produce.
Incorrect. Reporting only on exception turns the report into an escalation message. Progress reporting is periodic, and it reports against the plan and the schedule rather than against the defect database alone.
Test progress reports are produced once per test level after execution has finished and report the status of testing against the number of defects still open, because a status statement is meaningful only when every planned test case has been run.
Incorrect. That timing belongs to the test completion report. A progress report exists precisely because stakeholders need status information before execution has finished.
Test progress reports are produced at regular intervals but report progress against the development team's build schedule, because their purpose is to tell stakeholders which build reached the test environment rather than how testing itself is advancing.
Incorrect. The cadence is right but the reference point is not: a test progress report measures testing against the test plan and test schedule, with build information at most one item of context.
Test reporting summarises test information and communicates it to stakeholders. A test progress report is produced at regular intervals during a test activity, and it reports the current status of testing against the test plan and the agreed test schedule, including any deviation from them, so that stakeholders can take monitoring and control actions while those actions can still have an effect. The reporting frequency is agreed with the stakeholders and normally follows the reporting cycle of the project, for example weekly or once per iteration. The test completion report is a different artefact: it is produced at an agreed completion milestone and looks back at what was achieved.
Defect metrics for a car-sharing fleet platform can be grouped by source, by release, by test level, by priority and severity, by root cause and by status. Which statement correctly links one of these groupings to the management decision it best supports?
Grouping defects by root cause supports the selection of process improvement actions, because it shows which recurring causes produce the most defects and therefore where an upstream change would pay off best.
Correct. Root cause grouping is the grouping that identifies systematic weaknesses in the development or test process and so drives improvement actions rather than day-to-day control.
Grouping defects by status supports the decision on which additional test level to introduce in the next release, because the status of a defect report records the lifecycle phase in which the defect was originally introduced.
Incorrect. Status records where the defect report currently is in the defect workflow, not where the defect was introduced. Phase containment questions are answered by grouping by source or by test level.
Grouping defects by test level supports the decision on the order in which the open defect reports are triaged, because the test level in which a defect was found determines how urgently the fix has to be delivered.
Incorrect. Triage order is driven by priority, and to a lesser extent severity and risk. The test level a defect was found in is used for phase containment and test effectiveness analysis.
Grouping defects by severity supports the forecast of how much regression testing effort each developer will need, because the severity rating is a direct expression of the size of the code change required for the fix.
Incorrect. Severity expresses the impact of the failure on the system or the user, not the size of the fix. Effort for a fix and its regression testing has to be estimated separately.
Each grouping answers a different management question. Grouping by root cause points at what to change in the process upstream, which is why it is the natural input to process improvement. Grouping by status supports workflow and workload control, grouping by test level supports phase containment analysis, and grouping by severity and priority supports fix sequencing, not effort forecasting.
A test manager for a school meal ordering system is asked what test estimation actually produces and how the time-cost-quality triangle affects the answer. Which statement is correct?
Estimation produces effort, time and cost figures, and because time, cost and quality are interdependent, shortening the schedule without adding budget means the level of quality that testing can deliver has to be reconsidered.
Correct. Effort, duration and cost are the estimated quantities, and the time-cost-quality triangle states that a change to any one of the three constrains the others.
Estimation produces only a duration in calendar days, since effort and cost follow mechanically from the duration once the team size is fixed, and quality sits outside the triangle because it is owned by the development team.
Incorrect. Effort is the primary estimated quantity and duration is derived from it, not the other way round, and quality is one of the three corners of the triangle rather than something outside it.
Estimation produces effort in person hours for sequential projects and in story points for Agile projects, but the time-cost-quality triangle applies only where effort is measured in person hours, because story points already contain a quality allowance.
Incorrect. Both units are used to express estimated effort, and the triangle applies regardless of the unit. Story points express relative size, not a built-in quality allowance.
Estimation produces the cost of the test environment and of the test tools, while effort and time are planning inputs supplied by the project manager, and the triangle is used to decide which test techniques to select.
Incorrect. Effort, time and cost are all outputs of test estimation performed by the test manager, and the triangle supports trade-off decisions about scope and quality rather than the selection of test techniques.
Test estimation produces estimates of effort (expressed for example in person hours or story points), of the time needed and of the resulting cost. These three are linked through the time-cost-quality triangle: changing one of time, cost or quality forces a change in at least one of the others, so a fixed deadline with a fixed budget can only be met by adjusting the amount of testing and hence the achieved quality.
A test manager on a hospital staff rostering project is preparing the test estimate for the next release and is asked what that estimate has to cover. Which statement is correct?
The estimate must cover the effort of every activity in the test process, that is test planning, test analysis and design, test implementation, test execution and test completion, and it must include the work of setting up the test environment and preparing the test data.
Correct. Test estimation covers the effort of all test activities across the test process, and environment setup and test data preparation are part of that effort rather than something outside it.
The estimate must cover test execution alone, because test planning, analysis and design are management overhead already carried by the overall project estimate, while the test environment and the test data are delivered ready for use by the operations team.
Incorrect. Analysis, design, implementation and completion are test activities that consume test effort, and environment and test data work is normally a substantial part of the test estimate rather than a free delivery.
The estimate must cover the whole project, because a test estimate that leaves out development effort, requirements work and deployment cannot be compared with the budget that the project manager holds for the release as a whole, and the test manager is accountable for that comparison.
Incorrect. This confuses the test estimate with the project estimate. The test manager estimates the test effort, which is then one input into the project estimate owned by the project manager.
The estimate must cover the test activities but leave out test environment setup and test data preparation, because those are one-off infrastructure tasks whose cost belongs to the environment budget rather than to the estimated effort of the testing activities.
Incorrect. Environment setup and test data preparation are test activities and consume test effort, and leaving them out is one of the most common causes of an underestimated test effort.
Test estimation is an estimation of the effort that testing will involve, so it has to cover the activities that testing consists of: test planning, test analysis and design, test implementation, test execution and test completion. It explicitly includes the effort of setting up the test environment and of preparing test data, which are frequently underestimated because they are seen as infrastructure rather than as test work. Estimating the test effort is not the same as estimating the project: the test estimate is one input to the project estimate and does not cover development, requirements or deployment effort.
The defect workflow of a courier last-mile delivery app includes the statuses REJECTED, DEFERRED, RE-OPENED and CLARIFICATION in addition to the normal path. Which statement describes these statuses correctly?
REJECTED means the report is not accepted as a genuine defect, DEFERRED means the defect is accepted but its fix is postponed, RE-OPENED means retesting showed the failure is still present, and CLARIFICATION means the report has gone back to its author for missing information.
Correct. These are the standard meanings of the four statuses in the defect lifecycle described by the syllabus.
REJECTED means the fix was delivered but failed confirmation testing, DEFERRED means the report is waiting for a test environment, RE-OPENED means a duplicate was merged into an existing report, and CLARIFICATION means the severity rating is still under discussion in the triage meeting.
Incorrect. A fix that fails confirmation testing leads to RE-OPENED, and none of the other three statuses describes environment waits, duplicate merging or severity disputes.
REJECTED and DEFERRED are both terminal statuses in which no further action on the report is possible, while RE-OPENED and CLARIFICATION are informal annotations that a tester adds to the description rather than genuine statuses that the defect workflow recognises.
Incorrect. DEFERRED keeps the report alive for a later release, and RE-OPENED and CLARIFICATION are real workflow statuses that route the report back to defined owners.
REJECTED means the defect was found outside the agreed test scope, DEFERRED means the defect has low severity, RE-OPENED means the same defect was reported again in a later release, and CLARIFICATION means the root cause analysis of the report is still running.
Incorrect. These confuse status with other attributes: scope, severity, release and root cause analysis are recorded in separate fields and do not define these statuses.
Beyond the happy path a defect report can be REJECTED, meaning it is not considered a genuine defect, for example because it describes intended behaviour or a duplicate; DEFERRED, meaning it is accepted as a defect but the fix is postponed to a later release; RE-OPENED, meaning a report that had been closed or resolved has been found to still fail on retesting; and CLARIFICATION, meaning the report lacks information and has been sent back to the author for more detail.
A pension self-service programme has set up a defect management committee, holds regular triage meetings and has appointed a defect manager. Which statement describes these three elements correctly?
The committee is a cross-functional body that decides how reported defects are handled, the triage meeting is the recurring session in which it assigns priority and ownership, and the defect manager owns the process, the workflow and the quality of the defect data.
Correct. This matches the division of responsibilities in the syllabus: the committee decides, the triage meeting is where it works, and the defect manager is the custodian of the process itself.
The committee is made up of test managers only so that its decisions stay independent of the development organisation, the triage meeting decides the technical root cause of each report, and the defect manager assigns the developer who will implement each fix in the next sprint.
Incorrect. The committee is deliberately cross-functional, root cause is determined by analysis rather than voted on in triage, and fix assignment belongs to development management.
The committee reviews the severity ratings entered by testers and corrects them, the triage meeting is held once at the end of each test level to close the remaining reports, and the defect manager acts as the deputy of the test manager during test execution.
Incorrect. Severity is a technical assessment made when the report is raised, triage is a recurring activity throughout the project, and the defect manager is a process owner rather than a deputy test manager.
The committee is responsible for the test completion report, the triage meeting is a discussion between the tester and the developer who received the report, and the defect manager maintains the defect management tool configuration for the tool administrator.
Incorrect. The completion report belongs to the test manager, triage is a cross-functional decision forum rather than a bilateral discussion, and the defect manager's remit is the process, not tool administration.
The defect management committee is a cross-functional body of representatives of the stakeholders that decides how defect reports are handled, in particular priority and target release. The triage meeting is the recurring working session in which the committee reviews new and disputed reports and assigns priority and ownership. The defect manager owns the defect management process itself: the workflow, the data quality of the reports and the reporting on defect information.
A test manager for a sports timing system is designing the defect report template and has to decide which fields testers fill in manually, which the tool generates, and how many fields the template should have in total. Which statement reflects good practice?
Identifier, creation date and author should be generated by the tool while the tester supplies the observational and judgement fields, and the template should be limited to data that will genuinely be used, because every mandatory field costs effort on every report raised.
Correct. It separates tool-generated from manually entered fields and applies the principle of collecting only data that will be used.
As many fields as the tool supports should be made mandatory from the start, because data that is not captured at the moment of discovery can no longer be reconstructed later, and unused fields cost the team nothing once the template has been configured in the defect management tool.
Incorrect. Unused mandatory fields are not free: they cost tester effort on every report and tend to be filled in carelessly, which degrades the quality of the data that is actually used.
Severity and priority should be left out of the tester's part of the template and derived by the tool from the test case that failed, so that the two ratings stay objective and the tester only has to describe the behaviour that was actually observed in the session.
Incorrect. Severity is a judgement about the impact of the failure and priority is a business decision; neither can be derived automatically from which test case failed.
Steps to reproduce and the software version should be optional fields so that reports can be raised quickly during exploratory sessions, and the missing details can then be added by the developer who picks the report up at the next triage meeting of the committee.
Incorrect. Reproduction steps and version are exactly the information the developer needs and cannot supply. Making them optional produces reports that cost more to process than they save at entry.
Some fields are generated by the defect management tool itself, such as the unique identifier, the date of creation and the identity of the author, so asking the tester to type them wastes effort and introduces errors. The tester supplies the fields that require judgement and observation: a descriptive summary, the steps to reproduce, expected and actual results, severity and priority, and the environment and version. The guiding principle is that only data which will actually be used should be collected, because every additional mandatory field costs effort on every report.
During system testing of a hotel revenue management product, a tester runs an…
During system testing of a hotel revenue management product, a tester runs an overnight rate-recalculation scenario and observes that the recalculated rate for one property is written twice, producing a duplicated audit entry. The tester records the anomaly and raises a defect report with the steps taken, the software version, the environment identifier and log extracts. The report reaches the developer, who runs the same steps eleven times on the same build and on the same environment and never sees the duplicate entry. The developer's own log extracts differ from the tester's, and the tester, when asked, cannot say whether an overnight batch job was running at the same time. The developer wants to close the report. The defect workflow in use includes the statuses NEW, ASSIGNED, CLARIFICATION, REJECTED, DEFERRED, RESOLVED, RE-OPENED and CLOSED. Applying the failure to anomaly to defect report chain and the defect lifecycle, what is the correct next status transition for this report?
Move the report to CLARIFICATION so that the author can add the missing context about concurrent batch activity and the matching log window, because the report is not yet analysable rather than invalid.
Correct. Non-reproducibility caused by missing contextual information is what the CLARIFICATION status exists for; the report returns to the author before any decision about its validity is taken.
Move the report to REJECTED because eleven unsuccessful reproduction attempts on the same build and environment establish that the reported behaviour is not a defect in the product but an observation error by the tester.
Incorrect. Failure to reproduce is not evidence that the behaviour is not a defect, particularly where a plausible unexamined variable exists, and rejection is a triage decision rather than a developer decision.
Move the report to DEFERRED so that it stays in the backlog for the next release, because the duplicated audit entry has no direct effect on the rates offered to guests and can be investigated when the team has more capacity.
Incorrect. DEFERRED means an accepted defect whose fix is postponed. Here the defect has not yet been analysed at all, so deferring hides an unanalysed report instead of resolving the information gap.
Move the report to RE-OPENED and assign it to a second developer, because a report that one developer could not reproduce has to be independently confirmed by another before any further status change is allowed.
Incorrect. RE-OPENED applies only to a report that had previously been resolved or closed and then failed again on retesting. This report has never been resolved.
A failure was genuinely observed and reported, so the report is not invalid; but the developer cannot analyse it because information about the concurrent batch activity is missing. The lifecycle provides CLARIFICATION exactly for this: the report goes back to its author to supply the missing context, after which it can be analysed. Only if the information cannot be supplied and the failure remains unreproducible would a decision to reject or defer be justified, and that decision belongs to the triage forum, not to the individual developer.
A waste collection routing system assigns lorries to streets each morning…
A waste collection routing system assigns lorries to streets each morning. During system testing a tester finds that when a depot supervisor edits a route while a second supervisor has the same route open, the second save silently overwrites the first without any warning, and the lost changes cannot be recovered from the audit trail. The behaviour occurs only when two supervisors edit the same route within the same minute. Interviews with the operations department establish that depots are staffed by one supervisor per shift, that concurrent editing of one route has happened twice in the past three years, and that a workaround exists: the route can be re-entered manually in about ten minutes. The next release is a contractual delivery to the first three pilot depots, each of which has a single supervisor. Applying the definitions of severity and priority, how should this defect report be rated and why do the two ratings diverge?
High severity because the failure destroys data silently and irrecoverably, but lower priority because the pilot depots staff one supervisor per shift, the trigger has occurred twice in three years and a ten-minute workaround exists.
Correct. Severity follows the impact of the failure on the system and its data, while priority follows business urgency, which here is reduced by the rarity of the trigger and the available workaround.
Low severity and low priority, because the failure has been observed only twice in three years and none of the three pilot depots can currently produce the concurrent-editing condition, so the impact on the delivery is negligible.
Incorrect. Frequency of occurrence belongs to the priority judgement. Severity is assessed from what happens when the failure does occur, and silent unrecoverable data loss is severe whenever it occurs.
High severity and equally high priority, because a defect that causes unrecoverable data loss has to be fixed before a contractual delivery whatever the operational context of the pilot depots and their shift patterns turns out to be.
Incorrect. Tying priority mechanically to severity discards the business information about frequency, affected users and workaround, which is exactly what the separate priority attribute is for.
Low severity but high priority, because the technical fault is only a small locking omission in one save routine, while the reputational exposure at a contractual pilot delivery makes an immediate fix commercially urgent.
Incorrect. Severity reflects the impact of the failure, not the size or difficulty of the code change, so a small coding omission that destroys data is still a high-severity failure.
Severity describes the impact of the failure on the system and its data. Silent, unrecoverable data loss is a high-severity failure regardless of how seldom it happens. Priority describes the urgency of fixing it relative to other work, and it is a business decision that takes frequency, the affected user population and available workarounds into account. Because the pilot depots have one supervisor per shift, the trigger condition is nearly absent and a workaround exists, so priority is lower than severity. The divergence is expected and is precisely why the two attributes are recorded separately.
A digital radiology archive has completed four releases. In the last two…
A digital radiology archive has completed four releases. In the last two releases, 61 of 190 defect reports concern the DICOM import interface, and reading them shows a repeating pattern: eleven reports where an optional tag was absent and the importer aborted the whole study, nine where a vendor wrote a tag in a permitted but unusual encoding, seven where the study arrived split across two transfers, and six where a tag exceeded its documented length. The remaining reports in the cluster are one-off coding mistakes. The test manager wants to extend the existing defect taxonomy so that the recorded data will point to a concrete process improvement action rather than simply confirming that the importer is defect-prone. Which set of taxonomy categories should be introduced?
Categories for the kind of unhandled input variant, such as absent optional element, unusual but permitted encoding, fragmented transfer and out-of-range value length, because these distinguish the specification and test-basis gaps that a corrective action would target.
Correct. The categories mirror the causal pattern in the cluster and translate directly into an improvement action: tighten the interface contract and add a systematic set of input-variant test cases.
Categories for the affected architectural component, such as importer, parser, storage service and viewer, because knowing which component carries the most defects lets the team target its refactoring and its additional test effort where the risk is highest.
Incorrect. The cluster is already known to sit in the import interface, so component categories only restate what has been established and give no indication of what to change in the process.
Categories for the severity band of the failure, such as study rejected, study degraded and cosmetic display fault, because grouping by operational impact shows which of the import defects justify immediate effort and which can be deferred to a later release.
Incorrect. Severity is recorded on every report already and supports fix sequencing, not causal analysis. It cannot show why the same class of import problem keeps recurring.
Categories for the originating vendor system that produced the study, because tracing the reports back to the individual sending devices identifies the vendors whose conformance to the interface contract should be challenged before the next release of the archive is accepted into service.
Incorrect. This attributes the problem outside the team and does not identify an internal process weakness, whereas the reports show permitted variants that the importer should have handled.
A defect taxonomy is only useful if the categories it adds discriminate between causes that lead to different corrective actions. Here the cluster is dominated by input variants that the specification and the test basis did not cover: absent optional tags, permitted but unusual encodings, split transfers and over-length values. Categorising by the kind of unhandled input variant makes the pattern actionable, because it points at specification of the interface contract and at a systematic set of input-variant test cases. Categorising by component, by severity or by the vendor involved records data the team already has and does not identify what to change.
A test manager has to estimate the system test execution window for the next…
A test manager has to estimate the system test execution window for the next release of a car-sharing fleet platform using a metric-based technique. The measurement database holds the execution throughput of the three previous comparable releases: release 1 achieved 4.5 test cases per tester-day, release 2 achieved 6.0 and release 3, the most recent, achieved 7.5, giving an average of 6.0 test cases per tester-day. The release now being planned has 720 test cases to execute and six testers are available full time for execution, with no other duties. Management asks for a single duration in working days and wants to know what should be done about the spread in the historical data. Which answer applies the metric-based technique correctly?
20 working days, reported together with a range of roughly 16 to 27 days that follows from the observed per-release throughput, and with an investigation into what drives that spread rather than letting the average conceal it.
Correct. 720 / (6 x 6.0) = 20 days, and applying the extreme observed throughputs gives 16 and about 27 days, which is the range that should accompany the point estimate.
120 working days, because 720 test cases at an average of 6.0 test cases per day is 120 days of work, and the variation between releases can simply be absorbed by the contingency already held in the project budget.
Incorrect. 120 is the effort in tester-days, not the duration. Dividing by the six available testers converts effort into a duration of 20 working days.
27 working days, obtained from the lowest observed throughput of 4.5 test cases per tester-day, since a metric-based estimate is only defensible when it is built on the worst measurement in the data set, after which the spread needs no further analysis.
Incorrect. Taking only the worst observation replaces estimation with an undeclared contingency and still leaves the cause of the variation unexamined; the point estimate should use the representative figure and the range should be stated separately.
16 working days, using the 7.5 test cases per tester-day achieved in the most recent release, because the newest measurement reflects the current tooling and team composition best and the two older releases should be dropped from the data set.
Incorrect. A single observation is a weak basis, and discarding the earlier releases removes the very evidence that throughput is unstable, producing an optimistic estimate with no stated uncertainty.
The metric-based calculation is 720 test cases divided by 6 testers multiplied by 6.0 test cases per tester-day, which is 720 / 36 = 20 working days. The historical spread is real information, not noise: the same calculation at 4.5 gives 720 / 27 = 26.7, about 27 days, and at 7.5 gives 720 / 45 = 16 days. The estimate should therefore be communicated as 20 days with an explicit range of about 16 to 27 days, and the test manager should investigate what caused the throughput to differ so much between releases, because that driver may well be present again.
A ski resort is replatforming its lift pass system. The test manager currently…
A ski resort is replatforming its lift pass system. The test manager currently sends one identical weekly dashboard to two very different audiences. The executive steering group, made up of the sponsor, the finance director and the head of resort operations, meets monthly and decides on the go-live date and on the release of the remaining budget. The development lead runs the daily work of two development teams and decides where rework and additional testing effort go. The dashboard contains: test cases planned, executed, passed and failed; defect inflow and outflow by severity; requirements coverage against the release criteria; risk coverage; residual risk; code coverage per module; defect density per component; the blocked count; and effort spent against effort planned. The steering group complains that it cannot see what the numbers mean for the go-live decision, while the development lead says the dashboard tells him nothing about where to act. Analyse the situation and select TWO statements that describe the correct tailoring of the metric set. (Select TWO)
The steering group should receive requirements coverage and risk coverage against the release criteria, residual risk and effort spent against planned, because its decisions concern the go-live date and the budget and these metrics are already expressed in those terms.
Correct. These are the outcome-level metrics that map directly onto a go-live and budget decision, which is the only decision this audience takes.
The development lead should receive defect density per component, code coverage per module and defect inflow by severity, because those metrics locate the weak areas of the code and so support the rework and test-effort decisions he controls.
Correct. Component-level and code-level metrics are actionable for the person who allocates engineering effort, and they are the part of the dashboard he is missing.
Both audiences should continue to receive the identical dashboard, because two different metric sets create two competing versions of the truth, and the steering group needs the component-level detail in order to challenge the development lead's judgement.
Incorrect. Tailored reports drawn from one measurement base do not create competing truths, and a monthly steering group cannot use component-level detail to take a go-live decision.
The steering group should receive only the raw counts of planned, executed, passed and failed test cases, because absolute counts are neutral facts whereas coverage and residual risk figures embed test-manager judgement that could bias the budget decision.
Incorrect. Raw counts carry no information about what remains at risk, and removing the interpretation the test manager is responsible for leaves the steering group less able to decide, not more objective.
Reporting has to be tailored to the information needs of the audience and to the decisions that audience actually takes. The steering group decides on go-live and budget, so it needs outcome-level metrics expressed against the release criteria: requirements and risk coverage, residual risk and effort spent against plan. The development lead decides where engineering effort goes, so he needs component-level detail: defect density per component, code coverage per module and defect inflow by severity. Sending everything to everyone is not neutrality, it simply moves the interpretation work onto the reader.
System testing of a public library digital lending platform has run for four…
System testing of a public library digital lending platform has run for four weeks against a plan of 120 test case executions per week. The figures are: week 1, 118 executed, 71 percent of executed cases passed, 4 cases blocked; week 2, 96 executed, 79 percent passed, 17 blocked; week 3, 64 executed, 88 percent passed, 39 blocked; week 4, 41 executed, 94 percent passed, 58 blocked. The blocked cases are concentrated in the e-book licence enforcement and the inter-library reservation areas, both of which were rated high risk in the test plan. The product owner has seen the rising pass rate and is proposing to bring the release date forward by a week. Analyse the data and decide what is really happening and what the test manager should report.
Quality is not improving: the untested high-risk areas have dropped out of the executable pool, so the rising pass rate reflects a shrinking, biased denominator. Report pass rate against total planned cases, show the residual risk in the blocked areas and escalate the blockers.
Correct. The three signals read together show a biased sample, not an improving product, and the report has to restore the denominator and make the concentrated residual risk visible.
Quality is improving as expected late in a test level, because the defects found earlier have now been fixed and fewer failures occur, so the test manager should confirm that the exit criteria are within reach and support bringing the release date forward by one week.
Incorrect. A genuinely improving product would show a rising pass rate at stable or rising throughput. Here throughput has fallen by two thirds and the improvement is confined to the cases that could still be run.
Test execution productivity has fallen because the remaining test cases are longer and more complex, which also explains why they pass more often, so the test manager should report a revised throughput assumption and request two additional testers for the remaining weeks of the test level.
Incorrect. Longer cases would not raise the pass rate, and the data attributes the lost executions to blocking issues, so adding testers cannot recover throughput while the blockers remain.
The blocked count is an environment and test data problem that sits outside the quality signal, so the test manager should report the rising pass rate trend unchanged and track the blocked cases separately in the impediment log until the environment lets them be run.
Incorrect. Separating the blocked cases from the quality signal is what makes the pass rate misleading, because the blocked cases are exactly the high-risk ones whose result is still unknown.
The pass rate is computed over an ever smaller and increasingly biased denominator. As blocking issues remove the high-risk areas from the executable pool, only the low-risk, already stable cases remain to be run, so the pass rate rises while nothing about product quality improves. Falling throughput and a growing blocked count together show that the team is running out of executable material. The correct report expresses progress against the total planned execution rather than against the executed subset, shows the residual risk sitting in the blocked high-risk areas, and escalates the removal of the blockers instead of supporting an earlier release date.
An Agile team building an urban air-quality monitoring portal works in two-week…
An Agile team building an urban air-quality monitoring portal works in two-week sprints. Estimation happens inside a single 90-minute backlog refinement session attended by the whole cross-functional team, and the resulting estimates have to be available the same afternoon for sprint planning. The team has worked on this product for eleven sprints and knows the domain and the codebase well, but nobody in the team has been trained in building or decomposing work models, and there is no facilitator available to run several rounds of estimation spread over separate days. The test manager has decided to use an expert-based technique and now has to choose between Wideband Delphi and planning poker, and to justify the choice against the selection factors of time constraints and knowledge in modeling. Which justification is correct?
Planning poker, because the whole estimate must be produced inside one 90-minute session while Wideband Delphi depends on several separated rounds, and because the team has no modeling skill, whereas relative sizing draws on the product knowledge it already has.
Correct. Both selection factors point the same way: the available time cannot accommodate multi-round facilitated estimation, and planning poker does not require the modeling competence Wideband Delphi assumes.
Wideband Delphi, because its facilitated discussion rounds are the mechanism by which modeling knowledge is transferred to a team that lacks it, so the additional calendar time it consumes should be accepted as an investment in the team's estimation maturity.
Incorrect. The selection factor asks whether the modeling knowledge is present now, not whether the technique might build it later, and the sprint cadence leaves no calendar time for that investment.
Wideband Delphi, because a sprint commitment requires an absolute effort figure in person-hours which only Wideband Delphi produces, and the extra calendar days that its separated rounds need can be recovered by shortening the sprint review and the retrospective.
Incorrect. Agile teams commit on relative size and observed velocity rather than on absolute person-hours, and cutting the review and retrospective removes feedback loops rather than freeing genuine estimation time.
Planning poker, because it is a metric-based technique that computes the estimate from the team's recorded velocity over its eleven completed sprints, whereas Wideband Delphi is expert-based and would need historical modeling data that this team has not collected.
Incorrect. Planning poker is an expert-based technique; velocity is used afterwards to convert the agreed relative sizes into a forecast, and it is not what produces the estimate for each item.
Both candidates are expert-based, so the choice cannot be argued on the metric-based versus expert-based axis. Time constraints rule against Wideband Delphi, whose value comes from several facilitated rounds separated in time; the whole estimate here must be produced inside one 90-minute session. Knowledge in modeling also rules against it, because its rounds assume participants who can decompose the work into a model and argue about that model. Planning poker needs neither: it converges in minutes on relative sizes and draws on the domain and codebase knowledge the team already has from eleven sprints.
Halfway through a national parks campsite reservation programme the original…
Halfway through a national parks campsite reservation programme the original test effort estimate is overtaken by two events in the same month. First, the two most experienced testers leave the organisation and are replaced by three testers who are new to the domain and to the team's tooling. Second, an architectural decision moves the availability and pricing data from the relational store to a graph database that nobody in the test team has ever tested, and for which the existing test data generation utilities do not work. Test execution throughput and the defect detection profile from the first half of the project were the basis of the current estimate. The test manager has to identify which of the five groups of effort-influencing factors have actually been hit before producing a revised estimate. Analyse the situation and select TWO statements that correctly identify an affected factor group and its consequence for the re-estimate. (Select TWO)
The People group is affected, because losing two experienced testers and adding three newcomers lowers the average skill and experience and adds coaching load, so the throughput observed in the first half of the project can no longer be carried into the estimate.
Correct. Skills, experience and team composition sit in the People group, and their change invalidates the throughput assumption that the current estimate rests on.
The Product group is affected, because the move to a graph database changes the technology and the testability of the test object and breaks the test data generation utilities, which raises test design, test data and analysis effort for the affected areas.
Correct. The technology and testability of the test object belong to the Product group, and the loss of the test data utilities is a concrete effort increase that has to enter the revised estimate.
The Test results group is affected, since a change of technology is recorded there, and the re-estimate can therefore be produced by scaling the current defect count by the relative size of the newly introduced component.
Incorrect. The Test results group covers the defects actually found and the rework they cause, not the technology choice, and scaling a defect count by component size is not a re-estimation of test effort.
The Test context group is affected, since both the team change and the technology change alter the organisational context of the project, and the revised estimate should therefore be produced by the project manager rather than by the test manager.
Incorrect. Team skills belong to People and technology to Product, and responsibility for the test effort estimate stays with the test manager whatever the context change.
The People factor group covers the skills, experience and team composition of those doing the testing, and it is directly hit: two experienced testers have gone and three newcomers arrive with a coaching load, so the throughput assumption carried over from the first half of the project no longer holds. The Product factor group covers the characteristics of the test object, including the technology used and its testability, and it is hit by the move to a graph database that the team has never tested and for which the test data utilities no longer work. The re-estimate therefore cannot simply be scaled from the earlier measurements; the affected assumptions have to be replaced.
Chapter 3 · Managing the Team — 9 questions
A test team on a blood bank inventory programme has left behind the period in…
A test team on a blood bank inventory programme has left behind the period in which ownership of the test approach was openly argued over. Working agreements are now accepted by everyone, and the programme itself will close in three months. Which statement correctly characterises the norming and the adjourning phases of the Tuckman model and the test manager's task in each?
In norming the team accepts common rules and roles, so the test manager can shift from directing to coaching and delegate more; in adjourning the team is dissolved and the manager secures lessons learned and reusable testware and supports each member's next assignment.
Correct. Norming is characterised by accepted norms and reduced conflict, which permits a less directive leadership style, while adjourning is the dissolution phase where preserving experience and supporting the individuals are the manager's tasks.
In norming the test manager must still lead directively because the working rules have not yet been accepted by the members; in adjourning the same team is deliberately re-formed for the coming release so that the productivity it has built up is not lost again.
Incorrect. Directive leadership while rules are still being established belongs to forming, not norming, and adjourning means the team is disbanded rather than kept together for the next release.
In norming the team reaches its highest productivity and works largely autonomously, so the manager can withdraw entirely; in adjourning the roles inside the team are renegotiated after individual members have left, and the manager rebuilds the working agreements from the beginning.
Incorrect. Peak productivity and autonomous working describe performing; renegotiating roles after a change in membership sends the team back through earlier phases rather than describing adjourning.
In norming the members still work as separate specialists while the test manager defines the whole test process alone; in adjourning the team stops raising new defect reports and performs only confirmation testing until the release is finally handed over.
Incorrect. Working as unconnected individuals under a manager who defines everything is characteristic of forming, and adjourning describes the dissolution of the team, not a change in the type of testing performed.
In norming the team accepts shared rules, roles and working agreements, conflict subsides, and the test manager can move from directing to coaching, delegating more responsibility. Adjourning (dissolution) is the closing phase in which the team is broken up; the test manager secures the experience gained (lessons learned, documentation, reusable testware), gives feedback, and supports the members' transition to their next assignment.
A test manager on a telescope scheduling system notices that the cost-of-quality material in the Foundation Level syllabus and the material in the Advanced Level Test Management syllabus are not presented in the same way. How do the two approaches differ?
The Advanced Level follows Feigenbaum and classifies all quality-related spending into four cost categories, while the Foundation Level follows Boehm and describes how the cost of removing a defect grows the later it is detected in the lifecycle.
Correct. Feigenbaum provides a classification of quality costs into prevention, appraisal, internal failure and external failure; Boehm describes the escalation of defect removal cost across the lifecycle. The two answer different questions.
Both approaches use four categories with identical content, and Feigenbaum simply names the second category appraisal cost where Boehm names it cost of detection, so the difference between the two syllabi is purely terminological and has no practical effect.
Incorrect. The Boehm view is not a four-category classification at all, so the difference cannot be reduced to the naming of one category.
Feigenbaum splits quality costs by the organisational unit that carries them, such as test, development and support, while Boehm splits exactly the same costs by defect severity, so that critical defects are costed separately from minor ones.
Incorrect. Feigenbaum's split is by type of quality cost rather than by cost centre, and Boehm relates cost to the phase of detection, not to severity.
The Feigenbaum categories cover only the costs arising after release while the Boehm approach covers only the costs arising before it, so a test manager applies the one during the project and the other during the warranty period that follows.
Incorrect. The Feigenbaum categories explicitly include prevention and appraisal costs incurred during the project, and the Boehm view extends into production, where defect removal is at its most expensive.
CTAL-TM presents cost of quality following Feigenbaum, which classifies all quality-related expenditure into four categories: cost of defect prevention, appraisal cost, internal failure cost and external failure cost. CTFL v4.0 presents the Boehm view, which describes how the cost of removing a defect increases the later in the software lifecycle it is detected. The Feigenbaum view is a classification of spending; the Boehm view is a statement about cost escalation over time. They are complementary, not alternative names for the same scheme.
A test manager for a livestock traceability platform is preparing the argument for the cost-benefit relationship of testing to the steering committee. Why does the argument have to combine qualitative and quantitative benefits of testing?
Quantitative benefits such as avoided failure costs and reduced rework can be set against the test budget in figures, while qualitative benefits such as confidence in the release, reputation and better decision information cannot be monetised but still matter to stakeholders, so only both together give a complete cost-benefit picture.
Correct. The two kinds of benefit are complementary: the quantitative side makes the comparison with cost possible, the qualitative side captures value that no figure expresses.
Qualitative benefits are the ones the test team measures for itself and quantitative benefits are the ones the finance department measures, so the manager reports the quantitative figures upward to the steering committee and keeps the qualitative observations inside the team as input to its own process improvement work at the end of the release.
Incorrect. The distinction is whether a benefit can be expressed numerically, not which department records it, and qualitative benefits are precisely the ones that need to be presented to stakeholders.
Quantitative benefits are the defect counts and coverage figures produced during execution while qualitative benefits are the opinions gathered in retrospectives, and because opinions cannot be audited by the steering committee only the quantitative side can legitimately carry a cost-benefit argument.
Incorrect. Defect counts and coverage are test progress measures rather than benefits, and discarding the qualitative side removes exactly the value arguments, such as confidence and reputation, that a steering committee weighs.
Qualitative benefits are the benefits expected before the project starts and quantitative benefits are the ones actually measured after it has ended, so the manager needs both kinds in order to compare what was planned against the outcome that was finally achieved and to justify the next test budget.
Incorrect. The distinction is one of measurability, not of timing; benefits of both kinds can be estimated in advance and assessed afterwards.
Quantitative benefits can be expressed in figures and money: avoided external failure costs, reduced rework, fewer production incidents, shorter downtime. Qualitative benefits cannot be directly monetised: confidence in the release, protection of reputation, better information for release decisions, compliance evidence. A purely monetary argument omits effects that stakeholders genuinely value, while a purely qualitative argument cannot be set against the cost side of the equation, so both are needed.
In a disaster alerting programme, the testers report that their work feels…
In a disaster alerting programme, the testers report that their work feels pointless: severe findings are routinely overruled at the release meeting and the team is mentioned in management reports only when a failure reaches the public alerting channel. Which statement best describes the demotivating factors typical of test teams here and a suitable response from the test manager?
Testers are demotivated when their findings are overruled without justification and when testing is treated as less valuable than development; the manager responds by making test results visible to decision makers, having overruling decisions explained, and recognising test achievements alongside development achievements.
Correct. Ignored results, low perceived status and visibility only in the event of failure are the demotivators specific to test teams, and recognition plus transparent decision making are the manager's levers.
The dominant demotivator is that testers have too much freedom in selecting their test techniques, which leaves them uncertain about what is actually expected of them; the manager restores motivation by prescribing one mandatory technique for each test level and reviewing its use in every status report.
Incorrect. Professional discretion in choosing techniques is a source of motivation rather than of demotivation, and removing it narrows the tester's responsibility further.
The main demotivator is the defect management tool, because it makes each tester's individual output visible and comparable with that of colleagues; the manager should therefore stop reporting defect metrics to anyone outside the team and discuss them only in the team's own weekly meeting, where they cannot be compared between individuals.
Incorrect. The problem described is that findings have no effect, not that they are recorded; withdrawing test reporting would make the results even less visible to the decision makers.
Demotivation in test teams stems from the inherently repetitive nature of test execution tasks; the manager addresses it by rotating every team member through automation work in turn, so that manual execution no longer forms part of anyone's regular assignment during the release.
Incorrect. Repetition is a real issue and task variation helps, but it does not explain the situation described, and withdrawing manual execution altogether is neither feasible nor a response to findings being overruled.
Typical demotivating factors in test teams are that test results are ignored or overruled without explanation, that testing is treated as less valuable than development, that testers are perceived only as bearers of bad news, and that they are involved too late to influence anything. The test manager counters these by making test results visible to the decision makers, ensuring test achievements are recognised alongside development achievements, giving reasons when a finding is overruled, and involving testers early in the lifecycle.
A test manager in a forestry harvest planning organisation is choosing between…
A test manager in a forestry harvest planning organisation is choosing between training and education, self-study, peer learning, mentoring and coaching, and training on the job as development measures for the team. Which statement correctly assigns these measures to methodological competence and to social competence?
Training and education together with self-study are best suited to methodological competence such as test techniques and estimation, whereas social competence such as communication and conflict handling is developed mainly through peer learning, mentoring and coaching, and training on the job, where the tester interacts and receives feedback.
Correct. Codified method knowledge can be taught or read; social competence has to be practised in interaction with others, which is what peer learning, mentoring and on-the-job development provide.
Training on the job is the primary route to methodological competence because the methods are applied directly on the real project, while self-study is the primary route to social competence because working alone develops the self-discipline and resilience a tester needs when defending findings under pressure.
Incorrect. Self-discipline and resilience belong to personal competence rather than social competence, and self-study offers no interaction, which is what social competence requires.
Mentoring and coaching address professional and technical competence only, while training and education cover all four competence areas equally well, which is why formal classroom courses should be the default development measure for a test team whatever the competence gap identified in the skills matrix.
Incorrect. Mentoring and coaching are among the strongest measures for social and personal competence, and a classroom course cannot develop competences that only appear in interaction.
Peer learning develops methodological competence and formal training develops social competence, because a course is attended together with other participants from several organisations whereas peer learning is no more than a one-to-one exchange of method knowledge between two colleagues who already sit together on the same test team and share the same background.
Incorrect. This reverses the assignment: attending a course alongside others does not itself build social competence, and peer learning works precisely through the interaction between colleagues.
Methodological competence covers test techniques, estimation, planning and structured problem solving; it is conveyed efficiently by formal training and education and by self-study, because the content is codified and can be taught. Social competence covers communication, conflict handling, giving and receiving feedback and cooperation; it can only be developed in interaction, so peer learning, mentoring and coaching, and training on the job are the effective measures for it.
A regional blood bank inventory system has completed its first year in…
A regional blood bank inventory system has completed its first year in production. The quality manager has extracted the following figures for that year, each already assigned to one of the four cost-of-quality categories used in this syllabus. Defect prevention: requirements and design reviews, coding standards work and tester training, EUR 42,000. Appraisal: test analysis, design and execution effort plus test tool licences and test environment operation, EUR 96,000. Internal failure: correcting and retesting defects that were found before the release went live, EUR 68,000. External failure: emergency hotfixes, incident handling by the service desk, re-issuing incorrect stock reports and the contractual penalty paid to the health authority after a shortage went undetected, EUR 154,000. The board asks the test manager for the total cost of quality for the year and for an interpretation of how that total is distributed. Which answer is correct?
EUR 360,000 in total; conformance costs (prevention plus appraisal) are EUR 138,000 against EUR 222,000 of failure costs, so roughly two thirds of the spend goes on defects that were not prevented, which points to under-investment in prevention and appraisal.
Correct. All four categories are summed for the total, and the 138,000 to 222,000 split shows failure costs dominating, the classic indicator that the conformance side is underfunded.
EUR 206,000 in total; external failure costs are borne by the service and support budget after handover and are therefore reported outside the cost of quality, so the remaining split of EUR 138,000 against EUR 68,000 shows conformance costs dominating.
Incorrect. External failure cost is one of the four categories and belongs in the total regardless of which budget settles it; excluding it hides the most expensive consequence of poor quality.
EUR 138,000 in total; cost of quality means only the money deliberately invested in order to attain quality, so prevention and appraisal are counted while the EUR 222,000 of failure figures is reported separately as the cost of poor quality.
Incorrect. This is the conformance subtotal only. In the four-category model, cost of quality comprises both the conformance and the non-conformance costs.
EUR 222,000 in total; prevention and appraisal are planned project effort that would be spent on any release, so the cost of quality figure consists only of the internal and external failure costs of EUR 68,000 and EUR 154,000 that were actually caused by defects in the released system.
Incorrect. This is the failure subtotal. Prevention and appraisal are quality costs by definition and must be included, otherwise increased test effort would appear to be free.
Cost of quality is the sum of all four categories: 42,000 + 96,000 + 68,000 + 154,000 = EUR 360,000. Conformance costs, that is the money spent deliberately to achieve quality, are defect prevention plus appraisal = 42,000 + 96,000 = EUR 138,000. Non-conformance costs, the failure costs, are internal plus external = 68,000 + 154,000 = EUR 222,000, which is about 62 per cent of the total. The interpretation is that far more is being spent on dealing with defects than on preventing and detecting them, which typically indicates under-investment in prevention and appraisal.
A national driving licence registry has been in production for two years, and…
A national driving licence registry has been in production for two years, and the test manager is preparing the business case for the appraisal budget of the coming release. The quantitative side of the cost-benefit relationship of testing is expressed in this syllabus as Average Savings per Defect = Average External Failure Costs - (Average Appraisal Costs + Average Internal Failure Costs). Measurement over the last two years gives an average external failure cost of EUR 4,200 for a defect that reaches production, and an average internal failure cost of EUR 1,100 for correcting and retesting a defect that is found before release. The steering committee will approve the budget only if testing delivers an average saving of at least EUR 900 per defect. What is the highest average appraisal cost per defect the test manager may plan with, and at what appraisal cost per defect would the appraisal investment stop paying for itself altogether?
At most EUR 2,200 per defect, because 4,200 minus (2,200 plus 1,100) equals the required saving of 900 exactly, and the investment stops paying for itself at EUR 3,100 per defect, where the average saving falls to zero and any higher appraisal cost makes it negative.
Correct. Solving 4,200 - (x + 1,100) = 900 gives x = 2,200, and solving the same expression for a saving of zero gives x = 3,100.
At most EUR 3,100 per defect, because that is the point at which the average external failure cost is fully offset by the appraisal and internal failure costs together, and the investment stops paying for itself only above EUR 4,200, the full external failure cost of a defect.
Incorrect. EUR 3,100 yields a saving of zero, not the EUR 900 the committee requires, so the target saving must also be subtracted; and at EUR 4,200 the saving would already be minus 1,100.
At most EUR 3,300 per defect, because the required saving of EUR 900 is deducted from the average external failure cost of EUR 4,200, and the investment stops paying for itself at EUR 4,200 per defect, where appraisal costs as much as a defect reaching production does.
Incorrect. This omits the average internal failure cost of EUR 1,100, which the formula subtracts alongside the appraisal cost; at EUR 3,300 the actual saving is minus 200.
At most EUR 6,200 per defect, because the external failure cost, the internal failure cost and the required saving are added together to give the amount the appraisal activity may consume, and the investment stops paying for itself only above that combined figure.
Incorrect. The formula subtracts the appraisal and internal failure costs from the external failure cost; adding the three figures inverts the relationship and produces a ceiling higher than the cost of the defect itself.
Rearranging the formula for the target saving: Average Appraisal Costs = Average External Failure Costs - Average Internal Failure Costs - Average Savings per Defect = 4,200 - 1,100 - 900 = EUR 2,200. That is the highest average appraisal cost per defect at which the committee's target of EUR 900 is still met. The point at which the investment stops paying for itself is where the average saving falls to zero: 4,200 - (x + 1,100) = 0, so x = EUR 3,100; above that figure the appraisal and internal failure costs together exceed what a production defect would have cost, and the saving turns negative. The figures alone do not settle the decision: the qualitative side of the cost-benefit relationship, such as confidence in the release, protection of the registry's reputation and the evidence needed for regulatory scrutiny, is not captured in the per-defect arithmetic and must be argued alongside it.
A test team maintaining a hospital pharmacy stock system scores very highly on…
A test team maintaining a hospital pharmacy stock system scores very highly on professional competence: every member holds an advanced test design qualification, the automation suite is well engineered, and the product risk analysis is thorough and up to date. The syllabus distinguishes four areas of competence - professional, methodological, social and personal - and the test manager's assessment of the team places social competence far below the other three. The symptoms are consistent: developers say defect reports read as accusations and are frequently rejected as not reproducible without any dialogue; triage meetings end in argument rather than agreement; testers avoid raising quality concerns with the product owner; and two reviews of the prescription rules module were abandoned because the facilitator could not keep the discussion factual. Coverage figures and technique application, by contrast, remain exemplary. Analyse which test process activity is most damaged by this competence profile, and explain why a further technical training course will not repair it.
Test execution together with the defect reporting and triage that belong to it, because a technically correct defect report that antagonises the developer still does not lead to a fix; the remedy is mentoring, coaching and feedback in joint triage and review sessions, since communication and conflict handling can only be developed in interaction.
Correct. The deficient area is social competence, which is exercised when findings are communicated, defended and negotiated; a technique course develops professional competence, which is already strong.
Test design, because reviews that break down mean the specifications the testers work from stay unreliable; the remedy is an advanced course in specification-based and experience-based techniques so that the team can derive test cases without depending on agreement with the developers and compensate by broader coverage.
Incorrect. Design quality is demonstrably intact, and broader coverage does not help when the resulting findings are rejected in triage; the course would reinforce the strongest competence area instead of the weakest.
Test planning, because a team that avoids raising concerns with the product owner will accept schedules it cannot meet; the remedy is training in test estimation techniques such as three-point and expert-based estimation, so that the plan is defensible on its own figures and can be presented to the stakeholders without any negotiation being required.
Incorrect. Planning is affected, but a better estimate still has to be argued with stakeholders, so the measure depends on the very social competence that is missing; estimation training addresses methodological competence.
Test monitoring and control, because status reports from a team that argues with the developers will be distrusted by the stakeholders; the remedy is a tooling course on dashboards and automated status reporting so that the figures speak for themselves and are no longer coloured by the relationship.
Incorrect. Reporting is a symptom rather than the main damage, and automated figures do not carry the interpretation and negotiation that reporting requires; the measure again targets professional rather than social competence.
The four areas of competence map onto different parts of the test process. Professional and methodological competence drive test analysis and test design, and both are strong here, which is why coverage and technique application are unaffected. Social competence - communication, cooperation, conflict handling, giving and receiving feedback - is what the team relies on during test execution, where defects are reported, defended and negotiated with developers, and during the reviews it facilitates. A defect report that is technically accurate but phrased as an accusation is rejected, so the finding produces no correction: the value of the whole preceding analysis and design work is lost at the point of communication. A further technical course addresses professional competence, which is not the deficient area, and social competence cannot be transferred by instruction because it exists only in interaction. The appropriate measures are peer learning, mentoring and coaching, and training on the job - for example paired defect triage with feedback afterwards, review facilitation practised under a mentor, and agreed rules for wording defect reports.
A test team has for several years tested an on-premises desktop application for…
A test team has for several years tested an on-premises desktop application for veterinary practice administration, delivered in annual releases following a sequential lifecycle. The same six people are now moved to a cloud-based livestock traceability platform: a set of microservices exchanging events, integrated with national animal registers, delivered in two-week iterations with a continuous integration and delivery pipeline and a regulatory audit trail. All three context drivers have therefore changed at once: the application domain, the architecture and technology, and the software development lifecycle model. The test manager has two months before the first iteration starts and must decide which competences to acquire and in what order, and how the skills matrix should look once the transition is complete. Which analysis is the most appropriate?
Lifecycle competence first, because it governs how every activity is scheduled from the first iteration, then architecture and technology competence such as API and service testing and pipeline automation, while domain and regulatory competence is built incrementally with the product owner; the matrix afterwards shows target and actual level per person per competence, with at least two people covering each critical competence.
Correct. The way of working must be established before the technical skills are applied inside it, domain knowledge grows story by story, and redundant coverage in the matrix removes single points of knowledge.
Domain competence first and in full before any testing begins, since the test conditions are derived from the domain and the traceability regulation is unforgiving; technology and lifecycle competence can be picked up during execution once the first iterations are running, and the matrix afterwards records one nominated domain expert per subject area, so that the training effort is not duplicated across the team and each expert can be scheduled where the knowledge is needed.
Incorrect. In an iterative lifecycle domain knowledge is acquired story by story rather than up front, and nominating a single expert per area creates exactly the single point of knowledge a skills matrix is meant to expose.
Only the technology competence has to be acquired, because responsibility for the iterative lifecycle lies with the Scrum Master and the application domain is described in sufficient detail in the user stories themselves; the skills matrix is then replaced by a list of tool and platform certificates showing who holds which qualification and when it has to be renewed, which is easier to keep current than a competence assessment.
Incorrect. Testers have to work within the iterative lifecycle themselves, and a certificate list records qualifications rather than the actual competence levels against required levels that a skills matrix compares.
All three competences in parallel through a single combined classroom course covering cloud architecture, agile working and the traceability regulation, after which the skills matrix records every team member at the same target level, since a uniform competence level keeps iteration planning free of scheduling constraints and lets the test manager treat any team member as interchangeable with any other.
Incorrect. A single course cannot convey three unrelated competence areas to a working level, and levelling everyone identically discards the differentiation between actual and required competence that makes the matrix useful.
The lifecycle competence has to come first, because it determines how every test activity is scheduled and who does what from iteration one: continuous testing, a definition of done, testing inside the iteration rather than in a phase afterwards. Next comes the architecture and technology competence, because without service-level and API testing, pipeline-integrated automation and containerised environments, the team cannot deliver anything within a two-week iteration. The application domain competence, including the traceability regulation, is built continuously and incrementally with the product owner and the regulatory documents, since it grows with each story rather than being acquirable in advance. Afterwards the skills matrix should show, per team member, target and actual level for each competence, with at least two members covering every critical competence so that no competence rests on a single person.