NEW BOT Телеграм, страница - 376745437

Data Science & Machine Learning

@datasciencefun

72.1K subscribers

768 photos

1 video

68 files

677 links

Join this channel to learn data science, artificial intelligence and machine learning with funny quizzes, interesting projects and amazing resources for free

For collaborations: @love_data

Download Telegram

About

Blog

Apps

Platform

Data Science & Machine Learning

72.1K subscribers

Data Science & Machine Learning

machine-learning-cheat-sheet.pdf

👍7

3.56K views12:25

Data Science & Machine Learning

Python_Complete_cheatsheet.pdf

👍6

3.69K views05:44

Data Science & Machine Learning

Some interview questions related to Data science

1- what is difference between structured data and unstructured data.

2- what is multicollinearity.and how to remove them

3- which algorithms you use to find the most correlated features in the datasets.

4- define entropy

5- what is the workflow of principal component analysis

6- what are the applications of principal component analysis not with respect to dimensionality reduction

7- what is the Convolutional neural network. Explain me its working

👍8❤5

3.94K views05:44

Data Science & Machine Learning

Decision trees and Random forests?

Decision tree is a type of supervised learning algorithm (having a pre-defined target variable) that is mostly used in classification problems. It works for both categorical and continuous input and output variables. In this technique, we split the population or sample into two or more homogeneous sets (or sub-populations) based on most significant splitter / differentiator in input variables.

Random Forest is a versatile machine learning method capable of performing both regression and classification tasks. It also undertakes dimensional reduction methods, treats missing values, outlier values and other essential steps of data exploration, and does a fairly good job. It is a type of ensemble learning method, where a group of weak models combine to form a powerful model.

👍9

4.04K views08:43

Data Science & Machine Learning

😉5 Machine Learning Algorithms with Project Ideas

📉Linear Regression -> House Price Prediction
📈Logistic Regression -> Loan Default Prediction
🗞️ SVM -> News Classification
🏛️ KNN -> Breast Cancer Classification
🧮 Naive Bayes -> Text Classification

👍17❤8

4.22K views04:38

Data Science & Machine Learning

Data Science Interview Questions.pdf

👍5

4.12K views04:09

Data Science & Machine Learning

Supervised Learning Cheatsheet.pdf

👍6

4K views14:33

Data Science & Machine Learning

cheatsheet-machine-learning-tips-and-tricks.pdf

cheatsheet-unsupervised-learning.pdf

cheatsheet-supervised-learning.pdf

cheatsheet-deep-learning.pdf

👍4

4.22K views06:48

Data Science & Machine Learning

Top free Data Science resources

@datasciencefun

1. CS109 Data Science
http://cs109.github.io/2015/pages/videos.html

2. Data Science Essentials
https://www.edx.org/course/data-science-essentials

3. Learning From Data from California Institute of Technology
http://work.caltech.edu/telecourse

4. Mathematics for Machine Learning by University of California, Berkeley
https://gwthomas.github.io/docs/math4ml.pdf?fbclid=IwAR2UsBgZW9MRgS3nEo8Zh_ukUFnwtFeQS8Ek3OjGxZtDa7UxTYgIs_9pzSI

5. Foundations of Data Science by Avrim Blum, John Hopcroft, and Ravindran Kannan
https://www.cs.cornell.edu/jeh/book.pdf?fbclid=IwAR19tDrnNh8OxAU1S-tPklL1mqj-51J1EJUHmcHIu2y6yEv5ugrWmySI2WY

6. Python Data Science Handbook
https://jakevdp.github.io/PythonDataScienceHandbook/?fbclid=IwAR34IRk2_zZ0ht7-8w5rz13N6RP54PqjarQw1PTpbMqKnewcwRy0oJ-Q4aM

7. CS 221 ― Artificial Intelligence
https://stanford.edu/~shervine/teaching/cs-221/

8. Ten Lectures and Forty-Two Open Problems in the Mathematics of Data Science
https://ocw.mit.edu/courses/mathematics/18-s096-topics-in-mathematics-of-data-science-fall-2015/lecture-notes/MIT18_S096F15_TenLec.pdf

9. Python for Data Analysis by Boston University
https://www.bu.edu/tech/files/2017/09/Python-for-Data-Analysis.pptx

10. Data Mining bu University of Buffalo
https://cedar.buffalo.edu/~srihari/CSE626/index.html?fbclid=IwAR3XZ50uSZAb3u5BP1Qz68x13_xNEH8EdEBQC9tmGEp1BoxLNpZuBCtfMSE

Share the channel link with friends
http://t.me/datasciencefun

#freecourses

👍4😁2

4.55K viewsedited 15:32

Data Science & Machine Learning

The Data Science Design Manual.pdf

3.72K views15:37

Data Science & Machine Learning

Gant_Laborde_Learning_Tensorflow_js_Powerful_Machine_Learning_in.pdf

3.78K views16:49

Data Science & Machine Learning

Kubeflow_for_Machine_Learning_From_Lab_to_Production_by_Trevor_Grant.pdf

👍6

3.66K views08:28

Data Science & Machine Learning

Introduction to Machine Learning.pdf

👍5❤1

3.81K views08:29

Data Science & Machine Learning

Ultimate Guide to Data Cleaning.pdf

3.75K views09:01

Data Science & Machine Learning

3.38K views11:52

Data Science & Machine Learning

You are given a data set. The data set has missing values which spread along 1 standard deviation from the median. What percentage of data would remain unaffected? Why?

Answer: This question has enough hints for you to start thinking! Since, the data is spread across median, let’s assume it’s a normal distribution. We know, in a normal distribution, ~68% of the data lies in 1 standard deviation from mean (or mode, median), which leaves ~32% of the data unaffected. Therefore, ~32% of the data would remain unaffected by missing values.

🔥6👍3

3.04K views14:25

Data Science & Machine Learning

DATA SCIENCE INTERVIEW QUESTIONS
[PART-20]

1. What relationships exist between a logistic regression’s coefficient and the Odds Ratio?

The coefficients and the odds ratios then represent the effect of each independent variable controlling for all of the other independent variables in the model and each coefficient can be tested for significance.

2. What’s the relationship between Principal Component Analysis (PCA) and Linear & Quadratic Discriminant Analysis (LDA & QDA)

LDA focuses on finding a feature subspace that maximizes the separability between the groups. While Principal component analysis is an unsupervised Dimensionality reduction technique, it ignores the class label. PCA focuses on capturing the direction of maximum variation in the data set.The PC1 the first principal component formed by PCA will account for maximum variation in the data.PC2 does the second-best job in capturing maximum variation and so on.

The LD1 the first new axes created by Linear Discriminant Analysis will account for capturing most variation between the groups or categories and then comes LD2 and so on.

3. What’s the difference between logistic and linear regression? How do you avoid local minima?

Linear Regression is used to handle regression problems whereas Logistic regression is used to handle the classification problems.
Linear regression provides a continuous output but Logistic regression provides discreet output.
The purpose of Linear Regression is to find the best-fitted line while Logistic regression is one step ahead and fitting the line values to the sigmoid curve.
The method for calculating loss function in linear regression is the mean squared error whereas for logistic regression it is maximum likelihood estimation.
We can try to prevent our loss function from getting stuck in a local minima by providing a momentum value. So, it provides a basic impulse to the loss function in a specific direction and helps the function avoid narrow or small local minima. Use stochastic gradient descent.

4. Explain the difference between type 1 and type 2 errors.

Type 1 error is a false positive error that ‘claims’ that an incident has occurred when, in fact, nothing has occurred. The best example of a false positive error is a false fire alarm – the alarm starts ringing when there’s no fire. Contrary to this, a Type 2 error is a false negative error that ‘claims’ nothing has occurred when something has definitely happened. It would be a Type 2 error to tell a pregnant lady that she isn’t carrying a baby.

ENJOY LEARNING 👍👍

3.46K viewsedited 10:10

Data Science & Machine Learning

Thoughtful Machine Learning.pdf

3.32K views10:14