David Guggenheim
After a career in high-tech entrepreneurship, I returned to school and earned my Ph.D., from which my focus was on teaching the art and science of analytics to computer science and business students. However, my background demanded that my curriculum had to deliver real business value well beyond the end of the semester. My students should be imbued with the knowledge and tools that can shape their decision-making long after graduation and well into their careers, even after just a single semester.
But because of the technology and platforms available, my business students, including all of those at the Gies College Business, University of Illinois Urbana-Champaign, were left wanting. Only 5% went on to data science graduate studies (MIT, Stanford, Columbia, etc.), and the other 95% were not able to use the tools beyond the course because of complexity –we had to teach too much coding and not enough
to make better business decisions for accountants, financial analysts, production supervisors, marketing whizzes, and the rest. My students arrived with a semester of Python, and even the ‘A’ students didn’t remember a single line of code from the prior semester. So I had to spend another three weeks on coding out of a 16-week semester. This left me disheartened, but then technology moved forward as it always does.
A development occurred in the field of open source data mining platforms, one in which my vision for citizen data scientists could be realized – the advent of PyCaret 3.0 for predictive models, and in addition, the Explainable Boosting Machine (open source from Microsoft Research) for interpretable models. Before we introduce the open-source technologies, let me explain our philosophy for citizen data science through a series of tenets:
- Start with the business decision and work backward – what is nice to know about the technology, and what is needed to know to make better decisions?
- All time spent coding is a cost, and any time spent analyzing is a benefit.
- When preparing predictors for modeling, it’s all or nothing.
- Use the Dual-Experiment Data Mining Process flow, a CitizenAnalytics exclusive invention, to use template-driven analytics for comprehensive data mining and performance across almost all learning model families.
- Learn the economics of the learning model instead of the mathematical algorithms to make better decisions in the real world.
Our Vision
It took me a year to write the book, “Empowering the Citizen Data Scientist in Everyone,” and in addition to my knowledge of business analytics, it contains all of the data mining process flows in Jupyter notebooks hosted in Google Colab. The datasets, platforms, and mining notebooks are all in the cloud and open source so a student could use a Chromebook and conduct professional data science. This is business analytics as taught by the professor who helped Verizon predict customer churn in the early days of data science, and who solved a 50-year problem in data science with an undergrad.
The business students who complete this material will have the ability to perform predictive and interpretable models with errors within 1-1.5% of a technical data scientist, and because of its conceptual nature, the knowledge will stay with them for a long time. This closes the gap between subject matter experts and data scientists, which then allows a reallocation of resources within the firm such that the top-level data people will work on the more difficult problems instead of routine.
Our Mission
Transform Data Science Education with CitizenAnalytics LLC
Unveil the Power of Data with “Empowering the Citizen Data Scientist in Everyone”
CitizenAnalytics LLC introduces an innovative approach to data science education, as detailed in our latest publication: “Empowering the Citizen Data Scientist in Everyone”.
Revolutionize Data Science Learning
Innovative Curriculum: Based on the principles and insights from our groundbreaking book.
