1/165
Comprehensive vocabulary flashcards covering Business Analytics concepts, the SOAR model, data management, database structures, descriptive statistics, probability sampling, hypothesis testing, and regression analysis.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
How does increasing the amount of data available to address business questions help the business analyst?
Increasing data availability helps business analysts create more accurate forecasts, gain deeper insight, and eliminate uncertainties. Because business analysts serve as the middlemen between management and data scientists, they need as much information as possible to interpret it and relay it into real business decisions.
What is the risk of a business analyst having too LITTLE data available when addressing a business question?
Limited data increases uncertainty, produces less accurate forecasts, and makes it harder for the analyst to fully interpret business context before relaying recommendations to decision-makers.
Which individual is primarily responsible for interpreting data-related needs and serving as a liaison between data scientists and decision-makers?
Business Analyst
Data Scientist
Decision Maker
Technical Manager
Business Analyst
A data scientist builds a complex predictive model but doesn’t know which business question it should answer. Whose job is it to translate the business need into technical requirements?
The Business Analyst.
A hospital’s data team has access to more patient data than ever, but staff report feeling overwhelmed and unable to find the insights that matter. This scenario illustrates:
The analytics mindset
A business process
Data overload
Explanatory visualization
Data overload
A retail chain adds five new real-time data feeds to its dashboard, but analysts now take longer to find actionable trends than before. What concept does this represent?
Data overload.
What best describes the relationship between ‘Data’ and ‘Decision’ in the Information Value Chain?
Data is directly used to make decisions.
Data needs to be processed into information and contextualized to inform decisions.
Decisions are made independent of data.
Data is less important than context in decision-making processes.
Data needs to be processed into information and contextualized to inform decisions.
Put the Information Value Chain in the correct order: Decision, Data, Information, Knowledge.
Data → Information → Knowledge → Decision.
Which of the following activities most closely aligns with creating business value through standardized processes?
Cataloging customer complaints
Structured trainings for new employees
Recurring maintenance of company equipment
Assessment of employee performance
Structured trainings for new employees
Which of the following is the BEST example of a standardized business process (not just an activity)?
A documented, repeatable onboarding workflow that every new employee goes through in the same order — because it is coordinated, standardized, and involves both people and systems.
How could a set of Instagram posts about the quality of a new Ford Maverick Truck turn into knowledge capable of affecting a decision at Ford?
Social media posts serve as knowledge for Ford to make future decisions because the consumer reviews, comments, and posts about the Truck serve as information about what consumers think of the product. Ford gathers positive and negative feedback, forms a trend about what needs improvement, and uses that insight to adjust their process in developing future products.
A hotel chain collects thousands of unorganized TripAdvisor reviews about a new resort. Walk through how this raw data becomes a decision about renovating the resort’s pool area.
The unorganized reviews are raw data; filtering for pool-related comments and adding sentiment context turns it into information; recognizing a repeated complaint about pool cleanliness becomes knowledge; management using that knowledge to approve a renovation budget is the decision.
A shipping company records that 15% of packages arrived late last month. On its own, this number is just raw data. Once the analyst learns this occurred during a week of severe winter storms, the number gains meaning. This added layer is called ______ .
Context
A retailer sees "450" in a spreadsheet cell with no other information. Adding the label "average order value, Black Friday weekend" to that number is an example of adding ______ .
Context
True/False: All business functions use the same type of analytics.
False
Marketing Analytics and Operations Analytics rely on identical techniques and answer the same types of business questions.
False — each functional analytics area (marketing, accounting, financial, operations) is tailored to different business questions and techniques relevant to that function.
______ analytics is more likely to address the most efficient way to source a cigarette lighter from Shenzhen, China to a convenience store on Green Street in Champaign, IL.
Operations
______ analytics would be used to determine the most efficient warehouse layout to minimize employee walking distance during order picking.
Operations
An analyst evaluates revenue, costs, and profit margins to determine whether the company should invest in new equipment. Which type of analytics is primarily involved?
Marketing
Operations
Accounting
Financial
Financial
An analyst calculates the expected return on investment (ROI) for leasing versus purchasing a fleet of delivery trucks. Which type of analytics is this?
Financial analytics.
A beverage company wants to understand which customer segments respond best to a new product’s advertising campaign. This falls primarily under:
Accounting analytics
Marketing analytics
Financial analytics
Operations analytics
Marketing analytics
A sneaker brand wants to know which social media platform drives the most engagement for its Gen Z customer segment. Which type of analytics is this?
Marketing analytics.
Which of the following best describes how Financial Analytics and Operations Analytics differ in their focus within a company?
Financial Analytics focuses on the management and projection of financial performance, whereas Operations Analytics concentrates on improving process efficiency.
Operations Analytics prioritizes financial planning and investment decisioning, while Financial Analytics targets strategic corporate operations.
Financial Analytics and Operations Analytics both primarily deal with employee efficiency and turnover.
Both Financial Analytics and Operations Analytics are mostly concerned with customer facing interactions.
Financial Analytics focuses on the management and projection of financial performance, whereas Operations Analytics concentrates on improving process efficiency.
A manufacturing plant wants to reduce machine downtime — is this Financial or Operations analytics, and why?
Operations analytics, because it concerns process/equipment efficiency rather than financial performance or investment planning.
If a company was trying to track its actual daily sales as compared to its sales targets, would you use a dynamic or static report?
Dynamic
Static
Dynamic
A CFO wants a report that automatically refreshes every time new transaction data is entered into the system. Should this be built as a dynamic or static report?
Dynamic.
A business analyst presents a monthly PDF report summarizing sales results, but the report does not update automatically when new data is added. What type of report is this?
Predictive report
Static report
Prescriptive report
Dynamic visualization
Static report
An analyst emails a spreadsheet snapshot of Q3 results to leadership. The file will not reflect any Q4 changes unless manually resent. What type of report is this?
Static report.
Before collecting any data, an analyst at a theme park wants to know: “Is our ticket pricing data accurate, complete, and free of major gaps?” This concern belongs to the ______ step of the SOAR model.
Obtain the Data
An analyst asks, “Where can I find our customer loyalty program’s transaction history, and is it accessible?” This concern belongs to which SOAR step?
Obtain the Data
A company collects a large dataset of customer behavior but starts analyzing it without clearly defining the business problem. Which SOAR step was skipped?
Specify the Question
Obtain the Data
Report the Results
Analyze the Data
Specify the Question
A team jumps straight into building dashboards from raw sales data without ever agreeing on what business question the dashboard should answer. Which SOAR step did they skip?
Specify the Question.
Which phase of the SOAR Analytics Model involves deciding on the appropriate method for disseminating the findings of data analysis?
Specify the Question
Obtain the Data
Analyze the Data
Report the Results
Report the Results
Choosing between a live dashboard, a PDF summary, or an in-person presentation to share findings with executives falls under which SOAR step?
Report the Results.
An analyst at an insurance company is scanning a scatterplot of claim amounts vs. policyholder age, not yet sure what she’s looking for, when she notices a cluster of unusually high claims among a narrow age band. She flags it for further investigation but doesn’t yet know the cause. Which statement is most accurate?
This is explanatory visualization because she found something notable
This is exploratory visualization, and the finding itself doesn’t yet belong to any of the five analytics types until she investigates further
This is exploratory visualization, and the finding is automatically diagnostic analytics
This step belongs to “Report the Results” since she identified a pattern
This is exploratory visualization, and the finding itself doesn’t yet belong to any of the five analytics types until she investigates further
An analyst notices an unusual dip in weekly subscription cancellations every time the price is under 9.99, but hasn’t yet determined why. Is the analytics type diagnostic yet, or is this still exploratory visualization with an unclassified finding?
Still exploratory visualization — the finding doesn’t get classified as diagnostic (or any other analytics type) until the analyst investigates and determines the “why.”
A retailer’s exploratory analysis reveals a strange spike in returns every third Tuesday of the month. The analyst has not yet determined the cause. What is the analyst’s most defensible next move, based on the chapter’s framework?
Immediately build an explanatory chart for leadership announcing the cause
Treat the pattern as confirmed prescriptive analytics and recommend a policy change
Loop back to “Specify the Question,” framing a diagnostic question about why the spike occurs, before reporting anything
Ignore the finding since exploratory visualizations aren’t reliable
Loop back to “Specify the Question,” framing a diagnostic question about why the spike occurs, before reporting anything
An analyst exploring website traffic data notices unusually high bounce rates on mobile devices but doesn’t know why. What should the analyst do next according to the recursive nature of SOAR?
Loop back to “Specify the Question” and frame a new, more specific diagnostic question (e.g., “why is mobile bounce rate higher than desktop?”) before analyzing further or reporting.
While exploring earnings data for publicly traded companies, an analyst notices an unusual gap in the distribution right around zero. This unexpected pattern in the data is called an ______ .
Anomaly
A quality control analyst notices an unusual number of products failing right at the 3-year warranty mark, more than statistically expected. This unexpected pattern is called a(n) ______ .
Anomaly
An analyst at a pharmaceutical company plots thousands of clinical trial data points looking for unexpected patterns before deciding what story to tell. This is an example of:
Explanatory data visualization
Exploratory data visualization
The information value chain
Prescriptive analytics
Exploratory data visualization
A data analyst at a bank scans through thousands of transaction records purely to see if anything unusual jumps out, with no predetermined hypothesis. What is this an example of?
Exploratory data visualization.
A restaurant chain wants to know not just what happened to sales last quarter, but what action it should take next quarter given expected demand. This question calls for ______ analytics.
Prescriptive
An airline wants to know the optimal number of seats to overbook on each flight to maximize revenue while minimizing bumped passengers. This calls for ______ analytics.
Prescriptive
Think of your Major and answer the following question: Where would that data come from?
(Personal/major-specific — the data would come from internal movements such as company transactions, revenue, etc., and from social media/external sources showing how consumers feel about a product or service, plus other information data scientists collect.)
For your major, name one INTERNAL data source and one EXTERNAL data source you’d use to solve a real business problem.
(Personal — internal example: company transaction records/CRM; external example: U.S. Census Bureau demographic data or social media sentiment.)
How does the SOAR Analytics Model approach the process of handling business data?
Begin with specifying the question and follow with data acquisition, analysis, and reporting.
Start with data analysis, then collect data, formulate a question, and conclude with reporting.
Initiate by obtaining data, analyze it, pose relevant questions, and finish with reporting.
Collect data, report results, specify the question, and analyze data sequentially.
Begin with specifying the question and follow with data acquisition, analysis, and reporting.
Put these SOAR steps in the correct order: Analyze the Data, Report the Results, Specify the Question, Obtain the Data.
Specify the Question → Obtain the Data → Analyze the Data → Report the Results.
In a relational database, rows are called ______ and columns in a table are known as ______ .
records / fields
In a relational database, a single unique identifier for every record in a table is called a ______ , while an attribute referencing that identifier in another table is called a ______ .
primary key / foreign key
A poorly designed database stores customer names repeatedly in multiple tables. What problem is most likely to occur?
Faster queries
Data inconsistency
Better relationships
Reduced redundancy
Data inconsistency
A company stores a customer’s shipping address separately in its Sales table, Returns table, and Loyalty table, and one gets updated but the others don’t. What is this problem called?
Data inconsistency (caused by redundant data storage instead of normalized relational structure).
In the context of relational databases, which of the following best describes the relationship between primary keys and foreign keys?
Primary keys uniquely identify each record in a table, while foreign keys link that record to records in other tables.
Primary keys and foreign keys both serve to speed up data retrieval in a table.
Foreign keys uniquely identify each record in a table, while primary keys are used to link the tables.
Both primary keys and foreign keys are optional components not required for creating relationships between tables.
Primary keys uniquely identify each record in a table, while foreign keys link that record to records in other tables.
A Customers table has a CustomerID as its unique identifier. A Transactions table includes a CustomerID column referencing that same table. Which key is CustomerID in the Transactions table?
A foreign key (it references the primary key of another table).
Which attribute of the four V’s of Big Data is best described by the characteristic ‘speed of generation’?
Volume
Velocity
Variety
Veracity
Velocity
A stock exchange generates millions of trade transactions per second that must be processed in real time. Which of the four V’s does this primarily represent?
Velocity.
What distinguishes text data types, such as tweets and hashtags, from tabular data types structured into rows and columns?
Tabular data is not read by machines
Text data usually has structured format in databases
Text data is unstructured and often lacks the predefined model of tabular data
Text data is more space-efficient than tabular data
Text data is unstructured and often lacks the predefined model of tabular data
Why can’t a raw customer service call transcript be dropped directly into a standard spreadsheet row/column structure the way an invoice record can?
Because the transcript is unstructured text data with no predefined fields or format, unlike tabular data which has a fixed row/column structure.
When would the analyst prefer raw data to aggregated data?
Analysts prefer raw data because it’s unbiased, so they aren’t restricted by summaries and can take and interpret the data in alignment with their business goals.
What is one disadvantage of only having access to aggregated (already summarized) sales data instead of raw transaction-level data?
Aggregated data conceals underlying patterns, outliers, and granular detail the analyst might need — since it’s already been summarized/filtered according to someone else’s assumptions rather than the analyst’s own business goals.
Think of your Major and answer the following question: Where would that data come from?
The data would come from internal movements such as company transactions, revenue, etc. On social media by seeing how consumers feel about our product/service, and other information data scientists collect.
For your major, name one INTERNAL data source and one EXTERNAL data source you’d use to solve a real business problem.
(Personal — internal example: company transaction records/CRM; external example: U.S. Census Bureau demographic data or social media sentiment.)
A company records customer satisfaction as: 1=Very Dissatisfied, 5=Very Satisfied. What type of data is this MOST appropriately?
Ratio
Interval
Ordinal
Nominal
Ordinal
A restaurant asks customers to rate their meal as “Poor,” “Fair,” “Good,” or “Excellent.” What type of data is this?
Ordinal (the categories have a natural order/ranking but the “distance” between them isn’t precisely measurable).
Which type of data allows for ranking and can have methods like counting, grouping, and ranking applied for analysis?
Nominal data
Ordinal data
Interval data
Ratio data
Ordinal data
Class standing (Freshman, Sophomore, Junior, Senior) is an example of which data type, and why?
Ordinal data — the categories have a meaningful order/rank, but you can’t do arithmetic like subtraction or averaging on them meaningfully.
Which type of numerical data allows for the use of summing, averaging, and complex calculations such as multiplication, and is characterized by an equal interval between observations?
Interval Data
Ratio Data
Ordinal Data
Nominal Data
Interval Data
SAT scores range from 400–1600, with equal distances between points but no true “zero” score representing a total absence of ability. What type of data is this?
Interval data.
Which of the following actions should be taken to ensure data quality before analysis?
Check data types for consistency and known errors
Validate data completeness only
Clean data by removing valid entries
Perform final analysis to determine data relevance
Check data types for consistency and known errors
Before running any statistics on a newly imported spreadsheet, what is the FIRST data-quality action an analyst should take?
Check that each column’s data type is consistent and free of known/typo errors (e.g., text accidentally in a numeric field).
Which of the following statements accurately describes why data types must be consistently checked during the data preparation phase?
To ensure data quality by verifying the suitability of data types for analysis.
To prevent non-printable characters in data entries.
To remove whitespace from data.
To replace current values with statistically relevant estimates.
To ensure data quality by verifying the suitability of data types for analysis.
Why would an analyst specifically check whether a “ZIP code” column is stored as text rather than as a number?
Because ZIP codes should be treated as categorical/text data (leading zeros matter and math operations on them are meaningless), so checking data type consistency ensures the field is suitable for its intended analytical use.
Think of your Major and answer the following questions: a) What data would you need to solve a major related problem?
To solve a problem in my major (MIS), I would need internal operations data like customer service, sales transactions, etc. I also would need to look at how markets are related to our organization and other services.
For a specific problem in your field of study, name two DIFFERENT internal data sources you’d combine to solve it.
(Personal — e.g., combining sales transaction data with customer service ticket data to diagnose a product issue.)
A company collects customer browsing behavior without informing users. What is the main ethical issue?
Data size
Poor visualization
Lack of transparency and consent
Data redundancy
Lack of transparency and consent
A fitness app sells anonymized location data to third parties without disclosing this in its privacy policy. What is the primary ethical concern?
Lack of transparency and consent.
A grocery chain’s IT team is designing a database to link its “Transactions” table with its “Customers” table so analysts can see which customer made which purchase. Which of the following are true about how this link is built? (Select All That Apply)
A primary key uniquely identifies each record within its own table
A foreign key is used to create the relationship between the two tables
The tables must be merged into a single table before any analysis can occur
Common fields (like Customer_ID) allow the two tables to connect
A primary key uniquely identifies each record within its own table
A foreign key is used to create the relationship between the two tables
Common fields (like Customer_ID) allow the two tables to connect
An HR system links an “Employees” table to a “Departments” table. Which of the following would be TRUE? Employee ID is the primary key of the Employees table / DepartmentID appears as a foreign key in the Employees table / The two tables must be combined into one before HR can run any reports / A shared field like DepartmentID allows the tables to connect.
True: Employee ID is the primary key of Employees; DepartmentID as foreign key in Employees; a shared field like DepartmentID allows connection. False: the tables do NOT need to be merged before running reports.
True/False: Because the startup’s trade file exceeds several million rows, it should be stored and shared as an .xlsx file so it can be opened directly in Excel for analysis.
False
A company’s daily transaction log has grown to 2 million rows, exceeding Excel’s worksheet row limit. What file format should it be stored/shared in instead, and why?
.csv (or another flat/database format) — because .csv files aren’t bound by Excel’s row- capacity limits and are platform-independent, unlike .xlsx.
True/False: If 5% of the values in a “customer age” column are missing, the only acceptable response under the chapter’s guidance is to remove every record with a missing age value.
False
A dataset has 60%+ of its values missing in one column. Is it still appropriate to simply impute (fill in) those values with the column mean?
No — imputing more than roughly 60% of a column’s values distorts the true distribution and introduces significant model bias; at that point other strategies (or dropping the column) should be considered.
A company tracks employee engagement survey scores on a scale where zero doesn’t represent a true “absence” of engagement, only a low point on an arbitrary scale — this makes the score ______ data. By contrast, an employee’s exact tenure in years (where zero truly means no time employed) is ______ data.
interval / ratio
Fahrenheit temperature has an arbitrary zero point, while a company’s exact revenue in dollars has a true zero (meaning literally no revenue). Fahrenheit temperature is ______ data, while revenue in dollars is ______ data.
interval / ratio
An analyst imports a CSV file into Excel and notices missing formatting and formulas. Why?
CSV files do not store formatting or formulas
Excel removed them automatically
CSV is corrupted
Data was aggregated
CSV files do not store formatting or formulas
Why would a pivot table built in Excel not transfer over correctly if the underlying data is exported and re-imported as a .csv file?
Because .csv files store only plain text data separated by commas — they cannot retain formatting, formulas, or pivot table structures.
A dataset includes text reviews, images, and videos along with numeric sales data. What challenge does this create?
Data redundancy
Data variety and integration complexity
No storage issues
Lack of data
Data variety and integration complexity
A hospital’s patient records include structured lab results (numbers), scanned handwritten notes (images), and doctor dictations (audio/text). What Big Data challenge does combining these represent?
Data variety and integration complexity.
Which software tool is primarily used for running statistical tests on data as per the described functionalities?
Microsoft Excel
Tableau
Gretl
Alteryx
Gretl
Which tool would be most appropriate for running a formal regression analysis on econometric data: Tableau, Gretl, or Power BI?
Gretl (open-source econometric software built for statistical regression analysis).
Which of the following is NOT considered an effective practice for a company to ethically handle customer data?
Publishing customer data statistics periodically
Implementing data security protocols
Establishing penalties for data misuse
Encrypting credit card information
Publishing customer data statistics periodically
Which of these would violate ethical data governance: encrypting payment data, auditing third-party vendors, or publicly posting individual customer purchase histories without consent?
Publicly posting individual customer purchase histories without consent — this violates transparency/consent and privacy protection principles.
What does the median value represented in a box plot signify?
The highest value of the dataset
The middle value when the dataset is arranged in increasing order
The average value of the dataset
The point where no values exist above or below
The middle value when the dataset is arranged in increasing order
In a box plot, which line inside the box represents the point where 50% of observations fall above and 50% fall below?
The median line.
True/False: In a box plot, the box itself represents the interquartile range (Q1 to Q3), while the whiskers typically extend to the minimum and maximum values within the upper and lower fences.
True
The whiskers of a box plot always extend to the absolute minimum and maximum values in the dataset, with no exceptions.
False — whiskers extend only to the minimum/maximum values within the calculated fences; anything beyond the fences is plotted separately as an outlier.
An operations analyst is comparing delivery times (in hours) across two warehouses. Warehouse A has a range of 2–10 hours and a standard deviation of 1.2 hours. Warehouse B has a range of 1–15 hours and a standard deviation of 4.5 hours. Which statements are correct?
Warehouse B has more variability/dispersion in delivery times than Warehouse A
Range alone tells you the full shape of the distribution
If Warehouse B’s delivery time data is bell-shaped and roughly matches a normal distribution, roughly 68% of deliveries would fall within one standard deviation of the mean
Standard deviation for Warehouse A would be measured in “hours squared”
Warehouse B has more variability/dispersion in delivery times than Warehouse A
If Warehouse B’s delivery time data is bell-shaped and roughly matches a normal distribution, roughly 68% of deliveries would fall within one standard deviation of the mean
Store A has a sales range of $200–⊤800 with SD of $50. Store B has a range of $100–⊤900 with SD of $220. Which is true: Store B has more variability than Store A / Range alone confirms both stores have similar distribution shapes / If Store B is normally distributed, ~95% of sales fall within 2 SDs of the mean / Standard deviation is expressed in “dollars squared.”
True: Store B has more variability; if normally distributed, ~95% of sales fall within 2 SDs. False: range doesn’t confirm shape; SD is in the same unit as the data (dollars), not squared.
An HR analyst pulls hourly wage data for 500 employees. The median wage is $14/hour, but the mean wage is $19/hour due to a small number of highly paid specialists. This distribution shape is best described as:
Left-skewed
Symmetrical
Right-skewed
Right-skewed
A company’s employee tenure data has a median of 2 years but a mean of 1.1 years, pulled down by a large number of very new hires. What shape is this distribution?
Left-skewed (mean pulled below median by low-value observations/outliers).