Questions:
How much data? 0-100 million rows / year up to 1 billion rows / month up to 10 billion rows / day more?
How many sources? 1 (hint: this is unlikely) 2-5 (e.g. one, monolithic internal database + Twitter analytics + Google Analytics + marketing email platform) fewer than 20 (e.g. you’ve got several, unintegrated internal teams with their own platforms) fewer than 200 (most enterprises) more?
Which stack is more prevelant (including your web tier)? Java Python neither
Do you do “real” machine learning using full time employees who are specifically trained in the discipline?
Do you have a dedicated “data science” team who only does that and nothing else?
Do you need to see data update in literal realtime?
For infrastructure cost (i.e. compute time and storage), what does “expensive” mean to you? $500 / mo $5,000 / mo $1 million / year $10 million / year more?
Do you have internal, full time software engineers on staff?
How many? 1-3 fewer than 20 fewer than 100 more?
Do you expect non-engineers (i.e. analysts and business unit leaders) to know how to drag-and-drop create their own reports in something like Tableau?
Do you expect non-engineers (i.e. analysts and business unit leaders) to know how to write simple, raw SQL?
If you had to use time from an engineer to write code to finish the analytics part of a major project, would you be disappointed?
Across the largest pool of data you have, how long is too long to wait for an interactive query or a report to render for a user? 1 second 20 seconds 5 minutes 1 hour Overnight
If you had to throw 50% of your infrastructure away in 5 years and start over with something more complicated/expensive in order to accomodate your growth, would you consider that a failure?
Are you willing to avoid needing to start over in 5 years at the expense of it taking 3x as long or being 3x as expensive to get off the ground today?
Now go through and answer all the same questions again, except imagining your needs in 1 year, 5 years, and as far as you can realistically imagine into the future
The boxes
change detection download to data lake transform, stitch, rollup and summarize data marts machine learning and data science feedback into transaction systems business intelligence operations and observability quality assurance
Results:
Stitch Data + Aurora + dbt + Looker
Dagster + Snowflake + Apache Spark + dbt + Looker
Dagster + Snowflake + xarray/dask + JupyterHub + dbt + Looker
Kinesis/Kafka/Snowpipe
Delta Lake?
Prometheus + Graphana for all of the above
2023 © Pollen Analytics LLC. ALL Rights Reserved.