Community resourceWorksheet
OCR H446 2.2.2 Data mining: patterns from large data sets
Part 6 of 14 · H446 2.2.2 · Computational methods
Storing a great deal of data is not data mining, and H446 2.2.2 rewards students who can hold that line. A network of community venues supplies the booking records, and every finding here has to travel a chain from evidence to meaning to a decision the service could actually test.
Students will:
- distinguish collecting or storing large data sets from interrogating them to discover patterns
- build a chain from a finding to a proposed decision and an outcome measure
- state a relationship without claiming that one thing caused the other
- identify limitations arising from missing, biased or inconsistent records
- qualify a conclusion drawn from an unfamiliar data set and its context
Inside: 7 explanation cells, 1 multiple-choice question, 2 fill-in-the-blanks cells and 2 written answers. 20 marks, about 20 to 30 minutes.
Series: H446 2.2.2 · Computational methods, part 6 of 14.
Shared by Coding PathwayVerified teacher
- 12 cells
- About 30 minutes
- CC BY-SA 4.0
- Shared 31 Aug 2026
- Updated 3 Sept 2026
Preview
The whole resource, exactly as a class sees it. Answers and marking are held back.
Data mining: patterns from large data sets
A network of community venues records millions of bookings, cancellations, arrival times and event choices. Data mining can interrogate this large data set to discover useful information that is not obvious from individual records.
By the end, you will be able to
- define data mining accurately;
- distinguish collecting data from mining it;
- identify patterns, trends, relationships and anomalies;
- explain how a discovered pattern can support a decision.
Reactivate: a database stores and retrieves records; analysis asks questions across many records.
Collection is not mining
Data mining is the analysis or interrogation of a large data set to discover useful patterns, trends, relationships or anomalies. Recording each booking creates the data set; it does not itself discover anything.
The finding must be interpreted carefully. An association between late arrivals and one venue entrance does not prove the entrance caused the delay. Data quality, missing records and confounding factors affect how confidently the organisation should act.
Worked finding with a decision chain
Finding: bookings made within two days of an event have a much higher cancellation rate, especially for free evening sessions.
- Evidence: the pattern appears across many booking records, not one anecdote.
- Meaning: short-notice free bookings are associated with non-attendance.
- Action: test reminders or a waiting-list rule for that group.
- Measure: compare later cancellation and attendance rates.
- Caution: season, venue and event type may explain some of the relationship.
A strong answer connects discovery → contextual action → measurable consequence.
Which activity is data mining?
- AScanning one ticket barcode
- BAnalysing millions of booking records to find combinations linked with cancellation
- CEntering an attendee’s email address
- DBacking up the bookings table
- large
- patterns
- anomalies
- keyboards
Guided checkpoint: from result to use
Result: attendees who visit the access-information page are less likely to abandon a booking. Build a careful explanation:
- State the relationship without claiming causation.
- Suggest one decision the service could test.
- Name one outcome measure.
- Give one alternative explanation or data-quality concern.
Explain how the venue network could use this finding and one reason it should not assume the page itself caused the lower abandonment rate.
Use finding → action → outcome, then a caution.
Students type their answer here.
Practical complexity
Large data may arrive in inconsistent formats, contain missing or biased records, and include sensitive personal information. Preparing it can take substantial storage and processing. Many possible relationships may also produce coincidental patterns. The usefulness of mining therefore depends on question design, data quality, validation and lawful, proportionate use, not merely owning lots of data.
A school group analyses several years of learning-platform activity and finds that students who attempt more retrieval questions tend to achieve higher assessment scores. Explain the finding, a useful response and two limitations of the conclusion.
Do not claim that the relationship alone proves cause.
Students type their answer here.
Closed-book checkpoint
Complete each sentence from memory. There is no answer bank and correctness is held for teacher review.
Review your understanding
Before submitting, check that you can explain the central distinction in your own words, expose the intermediate state that supports your answer and apply the method in an unfamiliar context.