Data Processing and Information - Chapter 1 Notes
Data Processing and Information (Cambridge AS & A Level – Chapter 1)
Objective of the chapter: understand data vs information, the qualities of information, and how information is obtained and processed; explore data sources (direct and indirect) and the practical implications of data processing in ICT.
Data vs Information
Data: raw facts in various forms (characters, symbols, images, audio, etc.) that on their own have no meaning.
Information: details that have been given meaning, usually produced by processing data (often with computers).
Example (from the transcript): a set of postal codes or telephone dialing codes by itself is just data; when organized together it becomes meaningful information (e.g., Indian postal codes or UK dialing codes).
Data Processing (high-level idea)
Data processing is the manipulation of data so that it becomes information with meaning.
Data is stored as binary digits (zeros and ones).
Data can be saved and processed on fixed or removable media (e.g., pen drives).
Practical example mentioned: a source file (CSV) opened in a spreadsheet with formulas to produce meaning from data.
Quick recap formula-style reference used in physics in the talk: deviations, velocity, density, force (F), mass (M), acceleration (A) – illustrating how data could be represented as mathematical relationships in other subjects.
Direct Data vs Indirect Data; Sources
Direct data (original source data): collected specifically for a task or purpose; used for that purpose only. Examples:
Questionnaires
Interviews
Observations
Data logging (using a computer and a sensor; data is analyzed, saved, and the results are output as charts/graphs)
Indirect data (obtained from third parties for a different purpose): examples include data gathered for another purpose (e.g., electoral registers).
Advantages of direct data:
Reliability: you know exactly where it originated; you control data collection from a defined group.
Disadvantages of direct data:
Time and cost constraints; sample size may be small due to resources.
Advantages of indirect data:
Often larger data sets; less time and money required; can be used when direct collection isn’t feasible.
Disadvantages of indirect data:
May not be collected for the current purpose; potential misalignment with current needs.
Examples mentioned:
Direct data sources include questionnaires, interviews, observations, data logging.
Indirect data example includes the electoral register (for elections).
Data capture is intended to support decision-making; data quality hinges on source relevance and context.
Data Processing and Information Quality
Information quality depends on several factors. The key quality factors discussed:
Accuracy: information should be as error-free as possible; accuracy depends on the quality of the data before processing; errors in original data produce erroneous information. Verification and validation help improve accuracy.
Relevance: data must be collected for a clear purpose; irrelevant data wastes time and can mislead decisions.
Age (Timeliness): information should be up-to-date; data can become outdated and lead to incorrect conclusions (e.g., personal records like marital status that aren’t updated).
Levels of detail: information should have the right amount of detail; too much detail can obscure the key point.
Form and format: information should be succinct and free of extraneous data; useful for examination and decision-making.
Completeness: information should cover all relevant parts of the problem; gaps reduce usefulness and may lead to wrong actions.
Encryption (security of information, especially on the internet):
Encryption scrambles data so that only authorized parties can understand it; it does not prevent interception but makes data unreadable to interceptors.
Public vs private data: encryption uses keys; encryption is essential for online transactions and private communications.
Encryption can be used for stored data and for data in transit.
Encryption: How it Works
Core idea: encryption converts plaintext to ciphertext using an encryption key; the receiver decrypts with a corresponding decryption key.
Keys:
Symmetric (secret key) encryption: same key is used to encrypt and decrypt; faster but requires secure key exchange; risk if the key is intercepted.
Asymmetric (public-key) encryption: uses a pair of keys (public, private). Public key encrypts; private key decrypts; anyone can have the public key, but only the private key can decrypt.
Hybrid approach: entities often use asymmetric encryption to securely exchange a symmetric key, then use the symmetric key to encrypt the data (faster for large data).
Key length example: commonly used cipher lengths include 128-bit keys; the number of possible keys is , illustrating the vast key space.
Plaintext and ciphertext terms:
Plaintext: original readable data
Ciphertext: encrypted unreadable data
Encryption protocols and their roles:
IPsec: secure communication, authenticates computers and encrypts data packets; used for VPNs.
SSH (Secure Shell): secure remote login to perform operations on a remote computer; can be used for secure data transfer.
TLS (Transport Layer Security) / SSL (Secure Sockets Layer): secure web page transmission; TLS is the improved version of SSL; HTTPS uses TLS/SSL to secure websites.
HTTPS indicators: padlock icon in the browser; secure data transmission for web pages.
Purposes of SSL/TLS:
Enable encryption to protect data in transit
Authentication to ensure the communicating parties are who they claim to be
Integrity to ensure data is not altered during transmission
PCI DSS compliance for payment-card data handling
Increased customer trust when visiting secure sites
Digital certificates and Certificate Authorities (CAs):
A server presents a digital certificate to prove its identity; contains domain, organization, and the device for which the certificate is issued.
Certificate Authority (CA) issues certificates after performing checks; the CA signs the certificate with a digital signature and provides a public key infrastructure to verify authenticity.
If a CA is compromised, bogus certificates could be issued, enabling attackers to impersonate legitimate websites
Uses of encryption in everyday contexts:
Hard disk encryption: ensures data on storage is automatically decrypted when read by authorized software; protects data if the disk is accessed by others.
Email encryption: three components – (1) encrypt the connection to the email provider, (2) encrypt the messages themselves, (3) decrypt archived or saved messages when needed.
HTTPS websites: TLS/SSL protects data in transit between a user and a web server; HTTPS is the secure version of HTTP.
Practical considerations and risks:
Encryption improves privacy and security but can be a target for ransomware and other cyber threats; defenders may use firewalls to limit damage.
Use of encryption protocols has advantages (privacy, data protection) and potential disadvantages (complexity, performance costs, potential misuse by attackers).
Security implications: encryption helps protect sensitive data but does not make systems invulnerable; a compromised endpoint or weak certificates can undermine security.
Email and Web Security Details
Email encryption concepts:
Protects the content of email messages from interception; standard protection includes securing the connection, encrypting the message content, and decrypting saved messages.
HTTPS and TLS/SSL specifics:
Hypertext Transfer Protocol Secure (HTTPS) uses TLS/SSL to secure data transmitted over the web.
The presence of HTTPS usually indicates a secure connection, while a padlock icon signals that the site is secured.
Certificates and trust:
Digital certificates tie a domain to a public key and verified identity.
CAs perform checks before issuing certificates; compromised CAs undermine trust in the system.
Real-world examples cited: online payments, secure web pages, and remote server access via SSH.
Data Processing Types and Their Pros/Cons
Batch processing:
Processing occurs on data in batches (e.g., overnight) with little or no human intervention.
Common in payroll and other nightly/system-wide updates.
Advantages: cost efficiency, better use of resources when demand is low; disadvantages: results are not immediately up-to-date.
Master file vs transaction file:
Master file contains stable data (e.g., name, work number, department, hourly rate).
Transaction file contains data that changes over time (e.g., hours worked).
Processing combines master and transaction files to produce outputs (e.g., payroll); the transaction file is usually ordered the same way as the master file for efficient processing; validation checks are used to detect errors.
Online processing:
Direct interaction with the central computer; allows immediate processing of transactions (e.g., EFT, online payments, ATM transactions).
Uses direct access to locate records quickly rather than sequentially.
Online processing enables immediate feedback to users; data is processed as they are entered.
Real-time processing:
A subset of online processing where the response time must be immediate and with no delay; outputs depend on inputs in real time (e.g., sensor-controlled systems like greenhouse/air conditioning opening valves or turning on heaters).
Sequential vs direct access:
Sequential access scans records one by one until the required record is found.
Direct access goes straight to the required record without scanning preceding records.
Online vs batch processing trade-offs:
Batch: delayed processing but efficient for large data volumes; offline processing (e.g., payroll processing overnight).
Online: immediate processing; suitable for user-facing and time-critical tasks (e.g., EFT, want-to-pay systems).
Real-World Examples and Concepts from the Transcript
Greenhouse example (real-time): sensors detect conditions; system adjusts inputs (heater, fans, valves) immediately; illustrates real-time processing where output affects input.
Postal and dialing code examples illustrate data becoming information when organized meaningfully.
Examples of data formats and validation: questionnaires, DOB ranges, driving-age limits, and class-year constraints illustrate the practical use of data quality controls.
Validation and Verification (Data Accuracy Checks)
Verification:
Ensures data entry is correct, typically against the original source or during data transfer between storage media.
Methods discussed:
Visual checking by the person entering data.
Double entry: data is entered twice; computer compares; discrepancies are flagged for human correction.
Validation:
Ensures data values are reasonable for the intended use (i.e., data are plausible).
Validation checks (common types):
Present check: ensures required fields (primary keys) have a value; necessary to identify unique records.
Range check: validates numeric values lie within a defined range (e.g., max/min costs); uses lower/upper bounds.
Type check: ensures data type (numeric, alphanumeric, etc.) matches the field requirement.
Length check: ensures alphanumeric fields have the correct number of characters (not typically applied to numeric fields).
Format check: enforces a specific layout (e.g., social security numbers, dates).
Check digit: validates numeric data using a check digit computed from preceding digits (e.g., alternating weights 1 and 3; total used to derive the final check digit).
Lookup check: validates data against a limited set of valid entries.
Consistency check (integrity check): cross-field validation within the same record or across related records (e.g., class-year constraints with DOB).
Limit check: ensures values meet a minimum or maximum limit (e.g., minimum age for driving license).
Complementarity of verification and validation:
Verification can catch errors that validation cannot, and validation can catch errors verification may miss.
Other validation concepts mentioned:
Parity (parity) checks: ensure data integrity at the bit level; a parity bit tracks whether the number of 1s in a set of bits is even or odd.
Check sums and control totals:
Checksum: a value derived from the data file as a whole to verify transmission accuracy (byte-by-byte vs. file-wide, i.e., checksum vs. parity).
Hash total: similar concept used on larger files; the sum of selected numeric values (e.g., student IDs) is transmitted with the data; recipient recomputes and compares to detect changes.
Control total: calculated on numeric fields only (used similarly to verify data integrity).
Data processing and information flow:
Data starts in raw form, is processed, and is translated into readable formats such as diagrams, graphs, and reports.
Data processing methods detailed: batch processing, online processing, and real-time processing with their respective advantages.
Key Formulas and Numerical References (LaTeX)
Key length and key space: 128-bit key; number of possible keys =
Byte value range reference (conceptual): a byte encodes values from to (i.e., 256 possible values per byte).
Example year-range constraint (DOB example): DOB must be between 09/01/2004 and 08/31/2005 for a class entering in year 11.
Practical and Ethical/Practical Implications
Privacy and security: encryption protects personal data and online transactions but also introduces complexities (key management, potential misuse, ransomware risks).
Trust and compliance: SSL/TLS and PCI DSS standards help protect users and payments; trusted CAs are essential for maintaining secure communications; a compromised CA undermines trust.
Data quality and decision making: high-quality data (accurate, relevant, timely, complete) leads to better decisions; poor data quality yields poor outcomes.
Real-time and online systems: enable immediate decision-making and actions, but require robust infrastructure and reliable connectivity to avoid delays or failures.
Ethical considerations: collecting direct data imposes responsibility for privacy; indirect data usage must respect consent and purpose limitation.
Quick Reference: Terminology to Remember
Data vs Information: raw facts vs meaningfully processed data.
Direct data vs Indirect data: original collection vs third-party data.
Data processing: transforming data into information.
Encryption types: symmetric (secret key) vs asymmetric (public/private keys).
Protocols: IPsec, SSH, TLS/SSL, HTTPS.
Certifications: Digital certificates, Certificate Authorities (CAs).
Validation vs Verification: plausibility vs accuracy against source.
Data processing modes: Batch, Online, Real-time; Master vs Transaction files; Sequential vs Direct access.
Checks: Present, Range, Type, Length, Format, Check Digit, Lookup, Consistency, Limit; Parity, Checksum, Hash total, Control total.
Summary Takeaway
Data becomes information when meaning is applied through processing.
Quality information requires careful data collection (direct vs indirect), appropriate validation/verification, and suitable data processing methods.
Encryption and secure protocols (SSL/TLS, IPsec, SSH) protect data in transit and at rest, while certificates and CAs underpin trust in secure communications.
Different data processing approaches (batch, online, real-time) suit different applications, each with its own trade-offs in timeliness, cost, and complexity.
Awareness of both technical and ethical implications is essential for effective and responsible use of data in ICT.