Thinking in Pandas: How to Use the Python Data Analysis Library the Right Way By: Hannah Stepanek
![]()
Thinking in Pandas: How to Use the Python Data Analysis Library the Right Way By: Hannah Stepanek | Ebooks – Computer/Internet | PDF | 2.04 MiB
June 6th 2020 | ISBN: 148425838X | 186 pages
Author: Hannah Stepanek
Understand and implement big data analysis solutions in pandas with an emphasis on performance. This book strengthens your intuition for working with pandas, the Python data analysis library, by exploring its underlying implementation and data structures.
Thinking in Pandas introduces the topic of big data and demonstrates concepts by looking at exciting and impactful projects that pandas helped to solve. From there, you will learn to assess your own projects by size and type to see if pandas is the appropriate library for your needs. Author Hannah Stepanek explains how to load and normalize data in pandas efficiently, and reviews some of the most commonly used loaders and several of their most powerful options. You will then learn how to access and transform data efficiently, what methods to avoid, and when to employ more advanced performance techniques. You will also go over basic data access and munging in pandas and the intuitive dictionary syntax. Choosing the right DataFrame format, how to work with multi-level DataFrames, and how pandas might be improved upon in the future are also covered.
By the end of the book, you will have a solid understanding of how the pandas library works under the hood. Get ready to make confident decisions in your own projects by utilizing pandas—the right way.
What You Will Learn
Understand the underlying data structure of pandas and why it performs the way it does under certain circumstances
Discover how to use pandas to extract, transform, and load data correctly with an emphasis on performance
Choose the right data frame so that the data analysis is simple and efficient.
Improve performance of pandas operations with other Python libraries
Who This Book Is For
Software engineers with basic programming skills in Python keen on using pandas for a big data analysis project.Python software developers interested in big data
04
02
Section 1: Explains the underlying implementation of pandas and establish a foundation of understanding that will be built on in future sections.Chapter 1: An Introduction to Big Data Pandas
Chapter Goal: Introduce the reader to big data and some exciting big data problems that pandas has helped solve
No of pages – 7
Sub -Topics• A brief introduction to big data• Some examples of impactful problems that pandas has helped solve
• The limitations of pandas (aka when to use pandas and when to use a different library)
Chapter 2: How Pandas Works Under the Hood
Chapter Goal: Help the reader understand the data structures that pandas is built on
No of pages: 20
Sub – Topics
• A brief review of C vs Python performance
• A brief review of Numpy (the library that pandas is built on)
• What a single indexed data frame looks like underneath
• What a multi-indexed dataframe looks like underneath• What a multi-indexed multi-level-column data frame looks like underneath• How to choose the right data frame
Section 2: Help the reader understand how to load and normalize data in pandas efficiently.
Chapter 3: Loading and Normalizing Data in Pandas
Chapter Goal: Help the reader understand how to load and normalize data in pandas in a performant manner.
No of pages: 20
Sub – Topics
• A review of some of the pandas loaders (CSV, JSON, SQL, etc) and some of their most useful options
• Normalizing while loading data
• How to get the best performance when loading data
• Some gotchas to watch out for when loading data
Section 3: Help the reader understand how to analyze and manipulate data in pandas efficiently.
Chapter 4: Basic Data Access and Munging in Pandas
Chapter Goal: An introduction to accessing data in pandas for beginners.
No of pages: 5
Sub – Topics:
• Using dictionary-syntax, iloc, and loc to access data.
• Using merge, join, and concatenate to combine data.
Chapter 5: Reshaping Data
Chapter Goal: Help the reader understand when to employ certain data reshaping techniques.
No of pages: 20
Sub – Topics:
• Review of pivot and pivot table
• Review of transpose
• Review of stack and unstack
• Review of melt
Chapter 6: Apply: When to Use it, When Not to Use it, and How to Get the Best Performance
Chapter Goal: Aid the reader in understanding when to use Apply and how they can improve its performance.
No of pages: 10
Sub – Topics:
• Review some examples of when not to use Apply.
• Review some examples that warrant the use of Apply.
• The performance implications of using Apply
• How to implement a performant Apply
Chapter 7: Groupby
Chapter Goal: Help the reader understand how to use Groupby and what alternative approaches exist.
No of pages: 7
Sub – Topics:• How Groupby works underneath• How to get the best performance when using Groupby
• Groupby alternatives
Section 4: Help the reader understand advanced techniques for improving performance and how Pandas might be improved upon in the future.
Chapter 8: NumExpr: Performance Improvements Beyond Pandas
Chapter Goal: Help the reader understand how installing NumExpr can improve pandas performance.
No of pages: 10
Sub – Topics:
• A brief review of computer architecture with an emphasis on memory caching.
• How NumExpr improves performance of pandas operations
• Why eval and query are faster when NumExpr is installed
Chapter 9: The Future of Pandas
Chapter Goal: Help the reader understand where the pandas library is headed in terms of implementation and how it could be improved.
No of pages: 7
Sub – Topics:
• Where is pandas headed?
• What are problem areas/areas for improvement?
• Conclusion
13
02
Hannah Stepanek is a software developer with a passion for performance and is an open source advocate. She has over seven years of industry experience programming in Python and spent one and a half of those years implementing a data analysis project using pandas. She currently works at a small remote company called Hypothesis, building an annotation tool for annotating the web.
Hannah was born and raised in Corvallis, OR, and graduated from Oregon State University with a major in electrical computer engineering. She enjoys engaging with the software community, often giving talks at local meetups as well as larger conferences. In early 2019, she spoke at PyCon US about the pandas library and at OpenCon Cascadia about the benefits of open source software. In her spare time she enjoys riding her horse Sophie and playing board games.
18
02
Understand and implement big data analysis solutions in pandas with an emphasis on performance. This book strengthens your intuition for working with pandas, the Python data analysis library, by exploring its underlying implementation and data structures.
Thinking in Pandasintroduces the topic of big data and demonstrates concepts by looking at exciting and impactful projects that pandas helped to solve. From there, you will learn to assess your own projects by size and type to see if pandas is the appropriate library for your needs. Author Hannah Stepanek explains how to load and normalize data in pandas efficiently, and reviews some of the most commonly used loaders and several of their most powerful options. You will then learn how to access and transform data efficiently, what methods to avoid, and when to employ more advanced performance techniques. You will also go over basic data access and munging in pandas and the intuitive dictionary syntax. Choosing the right DataFrame format, how to work with multi-level DataFrames, and how pandas might be improved upon in the future are also covered.
By the end of the book, you will have a solid understanding of how the pandas library works under the hood. Get ready to make confident decisions in your own projects by utilizing pandas—the right way.
You will:
Understand the underlying data structure of pandas and why it performs the way it does under certain circumstances
Discover how to use pandas to extract, transform, and load data correctly with an emphasis on performance
Choose the right data frame so that the data analysis is simple and efficient.
Improve performance of pandas operations with other Python libraries
19
02
Establishes a foundation of understanding by exploring the underlying data structures that pandas is built onGuides the reader through architecting a pandas based solution by emphasizing performance
Uses simple, practical, and exploratory examples to empower the reader to recognize when to use a given pandas feature
06
05
300
01
https://covers.springernature.com/boo…
01
01
https://www.springer.com/9781484258385
01
Springer Nature Imprint
APR
Apress
01
01
APR
Apress
01
05
5235327
Apress
Berkeley, CA
US
02
20200821
2020
01
WORLD
08
0
gr
01
235
mm
02
155
mm
27
03
9781484258392
15
9781484258392
01
ISBN-13 hyphenated
978-1-4842-5839-2
DG
Apress
01
ROW
NP
10
20200821
02
Recommended Retail Price
01
BIC discount group code
ASPVBT
02
Product discount group
SPVBT
01
44.99
AUD
AU
20190612
01
Recommended Retail Price
01
BIC discount group code
ASPVBT
02
Product discount group
SPVBT
01
40.90
AUD
AU
20190612
02
Recommended Retail Price
01
BIC discount group code
ASPVBT
02
Product discount group
SPVBT
01
29.50
CHF
CH
R
2.5
20190612
02
Recommended Retail Price
01
BIC discount group code
ASPVBT
02
Product discount group
SPVBT
01
26.74
EUR
DE
R
7
20190612
02
Recommended Retail Price
01
BIC discount group code
ASPVBT
02
Product discount group
SPVBT
01
26.36
EUR
FR
20190612
02
Recommended Retail Price
01
BIC discount group code
ASPVBT
02
Product discount group
SPVBT
01
25.99
EUR
IT
20190612
02
Recommended Retail Price
01
BIC discount group code
ASPVBT
02
Product discount group
SPVBT
01
27.24
EUR
NL
20190612
02
Recommended Retail Price
01
BIC discount group code
ASPVBT
02
Product discount group
SPVBT
01
27.49
EUR
AT
R
10
20190612
01
Recommended Retail Price
01
BIC discount group code
ASPVBT
02
Product discount group
SPVBT
01
24.99
EUR
ROW
20190612
01
Recommended Retail Price
01
BIC discount group code
ASPVBT
02
Product discount group
SPVBT
01
22.99
GBP
GB
20190612
01
Recommended Retail Price
01
BIC discount group code
ASPVBT
02
Product discount group
SPVBT
01
1780.00
INR
IN
20190612
Apress
01
US
02
Y
NP
10
20200821
01
Recommended Retail Price
01
BIC discount group code
ADGNY14
02
Product discount group
DGNY14
01
24.99
EUR
ROW
20190612
01
Recommended Retail Price
01
BIC discount group code
ADGNY14
02
Product discount group
DGNY14
01
27.99
USD
US
20190612
Download Thinking in Pandas: How to Use the Python Data Analysis Library the Right Way By: Hannah Stepanek ( Size: 2.04 MiB ) :
Keywords: Thinking, Pandas, How, Use, the, Python, Data, Analysis, Library, the, Right, Way, Hannah, Stepanek
http://nitroflare.com/view/18AA119AA44CC20/eedcThinPaHotoUsthPyDaAnLithRiWaByHaSt.zip

