Data Managment Exam 1

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/187

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 9:30 PM on 9/30/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

188 Terms

1
New cards

Python typing: static or dynamic?

Dynamic. You don't declare variable types, and they're checked at runtime

2
New cards

How does Python mark code blocks?

Indentation plus a colon (Java uses { })

3
New cards

Python single-line vs multi-line comment

single-line. '''…''' or """…""" multi-line


4
New cards

Is Python compiled or interpreted?

Interpreted

5
New cards

Year Python was first released

1991

6
New cards

range(3, 10) produces what?

3, 4, 5, 6, 7, 8, 9 (stop is excluded)

7
New cards

range(17, 100, 2) means what?

Start at 17, stop before 100, step 2 (odd numbers)

8
New cards

Slice syntax

seq[start:end:step]. Start is inclusive, end is exclusive, default step is 1

9
New cards

s = "Python": s[-1], s[0:4], s[::2]

'n', 'Pyth', 'Pto'

10
New cards

How do you reverse a string or list?

[::-1]

11
New cards

my_list[-2::-1] on [10,20,30,40,50,60]

[50, 40, 30, 20, 10] (start at 2nd-last, go backwards)

12
New cards

my_list[:3] vs my_list[2:]

First 3 elements (indices 0-2) vs everything from index 2 onward

13
New cards

Are strings mutable?

No, immutable. Methods return a new string

14
New cards

"a,b,c".split(",") and "-".join(parts)

['a','b','c'] and 'a-b-c'

15
New cards

s.find("x") when "x" isn't found

Returns -1

16
New cards

"ab" * 3

'ababab'

17
New cards

List: mutable or immutable? Ordered?

Mutable, ordered

18
New cards

list.append(x) vs list.insert(i, x)

Adds to the end vs inserts at index i

19
New cards

list.sort() vs sorted(list)

Sorts in place (returns None) vs returns a new sorted list

20
New cards

b = a (a is a list): does it copy?

No. Both names point to the same list. Use a.copy()

21
New cards

How do you access element in row 2, column 1 of a list of lists?

matrix[2][1] (first index = sublist, second = element)

22
New cards

Access the red value of an RGB image pixel

image[row][column][color]

23
New cards

pop() vs remove(x)

pop() removes and returns the last item (or by index). remove(x) removes the first matching value

24
New cards
Tuple: mutable? Syntax?
Immutable, ordered. (1, 2, 3)
25
New cards
Why use a tuple?
Immutability. Can be used as dictionary keys
26
New cards
Syntax for a one-item tuple
(5,) (the trailing comma is required)
27
New cards
Dictionary: what is it?
Collection of key-value pairs in {}
28
New cards
Rules for dictionary keys
Unique and immutable (str, int, tuple OK. List not OK)
29
New cards
What can dictionary values be?
Anything, including lists and other dicts
30
New cards
d.get(key) vs d[key] for a missing key
get returns None (or a default). d[key] raises KeyError
31
New cards
d.keys(), d.values(), d.items()
All keys / all values / (key, value) pairs
32
New cards
Does "x" in d check keys or values?
Keys
33
New cards
Dictionary comprehension
{x: x**2 for x in range(5)}
34
New cards
Set: what is it?
Unordered collection of unique elements
35
New cards
How do you make an empty set?
set(). Not {} (that's an empty dict)
36
New cards
Remove duplicates from a list
set(my_list)
37
New cards
Set union, intersection, difference, symmetric difference symbols
Pipe |, &, -, ^
38
New cards
{1,2,3} & {3,4,5}
{3}
39
New cards
{1,2,3} ^ {3,4,5}
{1, 2, 4, 5} (in either, not both)
40
New cards
{1,2,3} - {3,4,5}
{1, 2}
41
New cards
List comprehension: squares 0 to 4
[x**2 for x in range(5)]
42
New cards
Why are dicts and sets fast for lookup?
Hash-based, so insert/delete/lookup are ~O(1) on average
43
New cards
Time to check x in list vs x in set
O(n) vs O(1)
44
New cards
Lists vs tuples performance
Tuples are immutable and faster for read-only use
45
New cards
File mode r
Read (default). Error if the file doesn't exist
46
New cards
File mode w
Write. Creates or overwrites (erases existing contents)
47
New cards
File mode a
Append to the end
48
New cards
File mode rb / wb
Read / write binary
49
New cards
Why use with open(...)?
File is automatically closed, even if an error occurs (prevents leaks)
50
New cards
read() vs readline() vs readlines()
Whole file as one string / one line / list of all lines
51
New cards
Best way to read a big file
Loop: for line in file: (one line at a time)
52
New cards
Does write() add a newline?
No. Add "\n" yourself
53
New cards
Text mode vs binary mode
Text decodes bytes to strings. Binary returns raw bytes
54
New cards
Modules for CSV and JSON
csv and json
55
New cards
Catch a missing file
try: ... except FileNotFoundError:
56
New cards
Why does error handling matter?
Prevents crashes, handles missing or bad files gracefully
57
New cards
What is journaling in a file system?
Tracks changes to reduce corruption after crashes (NTFS, ext4, APFS)
58
New cards
FAT32 max file size
4 GB. No journaling, very limited security. Used for USB drives and cards
59
New cards
Default file systems: Windows / Linux / macOS
NTFS / ext4 / APFS What is pathlib good for?
60
New cards
Path.cwd()
Current working directory
61
New cards
How do you join paths in pathlib?
The / operator: BASE / "data"
62
New cards
OUT.mkdir(parents=True, exist_ok=True)
Creates the folder and any missing parents. No error if it exists
63
New cards
DATA.glob("*.csv")
Finds all .csv files in DATA
64
New cards
stat.st_size
File size in bytes
65
New cards
st_mtime
Last modification time (content changed)
66
New cards
st_ctime
Windows: creation time. Unix: last metadata change
67
New cards
st_atime
Last access time
68
New cards
Convert a raw timestamp to readable
datetime.fromtimestamp(stat.st_mtime)
69
New cards
collections.Counter
Dict subclass for counting occurrences
70
New cards
Stream vs load a large file
Stream line by line: constant memory
71
New cards
csv.DictReader
Reads each row as a dict keyed by the header names
72
New cards
csv.DictWriter steps
Create writer with fieldnames, writeheader(), writerows(rows)
73
New cards
Sort by rating, ties broken by votes
sorted(rows, key=lambda r: (r["rating"], r["votes"]), reverse=True)
74
New cards
json.load() vs json.loads()
From a file vs from a string
75
New cards
json.dump() vs json.dumps()
To a file vs to a string
76
New cards
pd.read_csv(..., chunksize=N)
Reads a big CSV in chunks of N rows
77
New cards
usecols and dtype in read_csv
Load only needed columns / use smaller types to save memory
78
New cards
gzip.open(..., "rt")
Read a compressed file as text
79
New cards
What is JSONL?
One JSON object per line
80
New cards
What is NumPy?
Fundamental package for scientific computing, built around the N-dimensional array
81
New cards
ndarray: homogeneous?
Yes, all elements the same type
82
New cards
NumPy key attributes
shape, dtype, size
83
New cards
Ways to create arrays
np.array(), np.zeros(), np.ones(), np.arange()
84
New cards
What is broadcasting?
Operating on arrays of different shapes (e.g., array × scalar)
85
New cards
a * b for NumPy arrays
Element-wise product (not matrix multiplication)
86
New cards
Matrix multiplication / dot product
np.dot()
87
New cards
x.reshape(2,3) rule
Total number of elements must stay the same
88
New cards
Why is NumPy faster than lists?
Contiguous, C-style memory and typed values
89
New cards
Who created pandas, and when?
Wes McKinney, 2008, at AQR Capital
90
New cards
Where does the name "pandas" come from?
"Panel data" (an econometrics term)
91
New cards
What is pandas built on?
NumPy
92
New cards
Series vs DataFrame
1D labeled array vs 2D table (rows and columns)
93
New cards
axis=0 vs axis=1
Rows/index vs columns
94
New cards
df.sum(axis=1)
Sums across each row
95
New cards
df.shape vs df.size
(rows, columns) vs total number of cells
96
New cards
df.info(), df.head(), df.columns, df.describe()
Column types and non-nulls / first rows / column names / summary stats
97
New cards
df.set_index('PARTY')
Makes PARTY the row index
98
New cards
pd.concat([df1, df2])
Stacks rows (append)
99
New cards
pd.merge(df1, df2, on='key')
Joins on a common key column
100
New cards
Which pandas tools are for big data that doesn't fit in memory?
Dask, Vaex Inner join