Blog

Blog Categories

The EU Open Data Portal gives access to open data published by EU institutions, agencies and other bodies. Around 70 EU institutions, bodies or departments use the platform to make over 12,500 datasets available . I n this Jupyter Notebook we will retrieve data from open data portal " http://data.europa.eu/euodp/en/home ". The portal is based on the open source project CKAN. CKAN stands for Comprehensive Knowledge...
  Google became the main starting point for our online activities. The search engine processes about 40,000 searches every second or 3.5 billion searches per day. It records what people are interested in, what they worry about or where they want to travel. In a unique manner, the search engine captures trends in interests and behavior. Hidden racisms, sexual orientation or ad returns - check out the work by Seth...
Since the dawn of the digital age, the amount of data stored on servers has risen dramatically. With this increase, more and more firms are looking for talent that can handle their datasets and generate insights for business decisions. Google Trends shows that the global volume of the search term “Data Analyst” nearly tripled over the last 5 years. How does the increasing demand translate into earnings of data analysts in...
Das Erlernen neuer Programmiersprachen ist eine Investition ins Humankapital. Die Ermittlung des Return on Investment kann daher sehr aussagekräftig sein. Die Anforderungen für jede Branche und jeden spezifischen Job sind sehr spezifisch – eine verallgemeinerbare Antwort auf diese Frage zu finden, ist deshalb schwierig. Ein Ansatz könnte aber darin bestehen, die erforderlichen Softwarekenntnisse bei...
Companies use machine learning to improve their business decisions. Algorithms select ads, predict consumers’ interest or optimize the use of storage. However, few stories of machine learning applications for public policy are out there, even though public employees often make comparable decisions. Similar to the business examples, decisions by public employees often try to optimize the use of limited resources. Algorithms may assist...
Are you looking for real world data science problems to sharpen your skills? In this post, we introduce you to four platforms hosting data science competitions. Data science competitions can be a great way for gaining practical experience with real world data, and for boosting your motivation through the competitive environment they provide. Check them out, competitions are a lot of fun! Kaggle Kaggle is the best known platform...
Curious about neural networks and deep learning? This post will inspire you to get started in deep learning. Why are we witnessing this kind of build up for neural networks? It is because of their amazing applications. Some of their applications include image classification, face recognition, pattern recognition, automatic machine translation, and so on. So, let’s get started now. Machine Learning is a field of computer science that...
Seit Beginn des digitalen Zeitalters sind die Datenmengen von Unternehmen enorm angestiegen. Mit wachsenden Servern voller Daten hat die Nachfrage nach der Fähigkeit, 1) die erforderliche Infrastruktur aufzubauen und 2) die Datensätze zu verstehen und zu analysieren, stark zugenommen. Thomas Davenport und T. J. Patil haben bereits 2012 in der Harvard Business Review Data Scientist als " The Sexiest Job of the 21st Century "...
The open-source project R is among the leading tools for data science and machine learning tasks. Given its open-source framework, there are continuous contributions, and package libraries with new features pop up frequently. Currently, the CRAN package repository features 12,525 available packages. This post takes a look at the most popular and useful packages that have set the standards for solving data manipulation, visualization, and...
GBM is a highly popular prediction model among data scientists or as top Kaggler Owen Zhang describes it: "My confession: I (over)use GBM. When in doubt, use GBM." GradientBoostingClassifier from sklearn is a popular and user friendly application of Gradient Boosting in Python (another nice and even faster tool is xgboost). Apart from setting up the feature space and fitting the model, parameter tuning is a crucial task in...
AI Took My Job! Ken Jennings’ name is vaguely familiar to people, but why? Because his profound knowledge on all things trivial led to him being the unbeatable champion of a TV game show called Jeopardy! It also put him in the gunsights of IBM. They spent thousands of hours, invested millions of dollars, all just to build a machine named WATSON that could defeat him playing that TV-derived game. See how Ken deals with the...
Demand for professionals in data science and analytics is expected to rise significantly over the next years (cf.  this study  by IBM). In order to keep track of future job trends, we started the DataCareer Job Market Index (DJMI) in July 2017. We track job openings on the biggest online job board,  Indeed , in the fields of data science and analytics, data engineering, business intelligence, artificial intelligence and...