Thoughts on perfectionism and fiction inspired by “Carrie Soto Is Back”

I recently finished “Carrie Soto is Back” by Taylor Jenkins Reid, picking it up after a friend recommended it to me because I like tennis. The book is about a (former) tennis star’s decision to return to the game 5 years after retiring in an effort to protect or reclaim her record for most Grand Slams. I’ve never read anything by TJR before, but I kind of knew what to expect in terms of character archetypes and plot, and that’s pretty similar to what I got. This isn’t really a review of Carrie Soto, but here’s my review: it’s good enough that I read it quickly and never considered quitting, but I did not become a TJR fan. Now, onto the stuff I really want to talk about!

Fictionalizing tennis in the world of Carrie Soto

In the world of Carrie Soto, the roster of professional tennis players is almost entirely fictional, and that mostly includes former stars as well. Carrie Soto, the main character, holds the record for most Grand Slams won, but Nicki Chan equals her record at the 1994 US Open with a victory over Ingrid Cortez. Of course, you can easily verify that the 1994 US Open was not contested by women named Ingrid Cortez and Nicki Chan; it was Arantxa Sánchez Vicario and Steffi Graf, neither of whom exist in Carrie Soto (or at least, they’re not tennis stars). Other greats from that moment, like Monica Seles, Gabriela Sabatini, Chris Evert, and Martina Navratilova, are also missing. Not even Billie Jean King is mentioned! Interestingly, some actual players do get mentioned—I recall a reference to Bjorn Borg’s failed attempt to come out of retirement, for example.

But this is fiction, of course, so who cares? I was not bothered by this in the slightest. I was certainly curious to consider who actually won the tournaments discussed in the book, but it was very easy for me to suspend my disbelief/knowledge for the purposes of this book. Frankly if anybody told me they couldn’t get into the book because it was rewriting tennis history, I’d find them totally unreasonable. I’ve never read sports fiction, but this has to be the bare minimum for the genre (unless authors get around it by never saying what year it is).

The mechanics of professional tennis: so much diving

While the book is mostly NOT tennis matches, Jenkins Reid does have to show some tennis, and so she writes direct play-by-play as well as summarizing playing experiences. In both types of writing, she reveals that the mechanics (for lack of a better word) of professional tennis in the world of Carrie Soto are not true-to-life.

  1. Carrie serves unrealistically fast. It’s much too fast for the era of tennis she played in (in the 70s and 80s) at over 120 mph. In fact, every player whose game is described in Carrie Soto has a fast serve and hits the ball hard, which oversimplifies the diversity of game styles that tennis allows. But fine, whatever, I’m actually not really bothered by this. It’s a fictional universe, and we’re supposed to believe that Carrie is a special talent anyway. So let her serve over 120 mph.
  2. Women and men do not really play each other. Carrie and Bowe, another older player and major supporting character in the book, end up becoming training mates, and Carrie regularly beats Bowe when they play practice matches. In the book, both players are elite-level— they are reaching the second week of Grand Slam tournaments. It’s just not likely that a male player who is competitive at Grand Slam level professional tennis would regularly lose to a female player who is competitive at Grand Slam tennis. Elite-level men and women don’t even practice together that often (not never, but not often!). Elite-level women likely travel with a male hitting partner, but hitting partners are not professional tennis players. They are just very good tennis players. But you know what? I also was not bothered by this. Again, Carrie is supposed to be special, so I say she can beat Bowe all she wants.
    (Also, let me be clear: I am much more interested in women’s tennis than men’s tennis, and I don’t believe that direct competition between men and women should be a requirement for women to be professional athletes and be paid the same as men.)
  3. Diving is spectacular but rare; most players do not dive for balls. In Jenkins Reid’s play-by-plays, players are constantly diving. This is very unrealistic. Most diving that happens in real-world tennis happens at the net when a player is trying to make a volley, but Jenkins Reid had players diving in baseline rallies and even trying to dive to return serves! Maybe Gaël Monfils would do something like this, but it’s also important to point out that whenever Monfils (or any other player) dives for a ball, it’s a big deal! So it could not be happening with such regularity. This is another thing that I was not bothered by (well, mostly!). Maybe tennis in the Carrie Soto universe is just so high stakes and exciting that players are willing to dive.

These details are not as easy to accept as the roster of fake players, but I can still accept them as part of the Carrie Soto universe. I mean, these are all statistical details anyway, and so it’s easy to imagine that the distribution of serve speeds, talent across gender, and umm… needing to dive… could be different in a fictional universe. Did it take me out of the story? Sometimes, but I can think something is kind of stupid without rejecting it as impossible. And I did!

Unbreakable rules in tennis (and geography)

What I could not accept—and what I was constantly taking pictures of and texting to my tennis-playing friends—was Jenkins Reid misunderstanding aspects of how a tennis match unfolds (scoring, changeovers) and not understanding or properly using relatively basic tennis jargon.

How tennis proceeds: set scoring and changeovers

There were three instances I recorded where moments in a Grand Slam tennis match are described, but those moments are not actually possible. I’m gonna kind of go into detail here for the benefit of non-tennis fans.


“I am sitting in my hotel room, watching Nicki play Andressa Machado. She has one set behind her; it’s 7-6 in the second. Machado is serving, and Nicki is running all over the court, making every shot.”

“7-6 in the second” only means one thing: the second set is over. That’s the score of a set when it has gone to a tiebreaker and the tiebreaker is completed. It’s pretty clear that Jenkins Reid did not intend to communicate that the second set was over— so it seems that she does not know that “It’s 7-6 in the second” is not something anybody would ever say about a modern tennis match. It could have just been 6-5 in the second set and then there wouldn’t be a problem.


“We are now tied 6-6 in the second set. She is serving for the set. She sends three kick serves in a row, and each one bounces differently. It knocks me out of my flow. The second set is hers.”

Similar to the last example, we’re dealing with a tiebreak set here. When a non-final set reaches a score of 6-6, a tiebreaker begins. When a tiebreaker is reached, there is no longer somebody “serving for the set” because players alternate serves during the tiebreaker. And no player would serve three times in a row in a tiebreaker (unless one of the serves was a fault or a let) because each player only serves two points in a row in a tiebreaker. All of this suggests that Jenkins Reid doesn’t know that at 6-6 in the second set, the players would go to a tiebreaker.


“It’s the final set, 4-4. … During the changeover, I sit down to drink my water. I breathe in deeply and close my eyes. I have to rethink my strategy here.”

Changeovers—where players get to sit down before continuing play on the other side of the court—occur at two times in tennis matches. First, there’s a changeover in between every set, no matter the score. Second, there is a changeover only after odd-numbered games. (After the first game of a set, players are technically supposed to switch sides without taking a break, but in practice, most professional tennis players will take a quick changeover after the first game of the set as well). A score of 4-4 is an even number of games, so there would not be a changeover at 4-4. Again, this is sort of a pointless detail to mess up; she could have just said it was 3-4, and the story would not be affected at all. But it demonstrates that she (and her editor) don’t understand how a tennis match progresses.

Games and matches

Since game is commonly used to refer to a complete contest (athletic or leisure), it would be nice if a complete tennis contest was also called a game. But game has a more specialized use in tennis, and so a different word is used for a complete contest: a match. For clarity, this is how tennis scoring proceeds: win the most points to win the game, win the most games to win the set, win the most sets to win the match. That’s where the saying “Game, set, match!” comes from: the player has won a point which won them the game, which won them the set, which won them the match.

Several times in Carrie Soto is Back, a character who really should know better (e.g., a professional player or renowned coach) would refer to a completed contest as a game. For example, Carrie’s father and coach Javier might have complimented her performance in the match by saying that she played a great “game.” The trouble with this inaccuracy is that there is a group of people who say things like this: people who don’t know very much about tennis. When these characters improperly use the term game when talking about a tennis match, they don’t just sound wrong; they sound like amateurs!

Break points

There also seems to be some confusion around the term “break point.” For newcomers, here’s a description of a break point. When you are the receiving (i.e., non-serving player) and you win enough points to be one point away from winning that game, you have now reached a break point. Essentially, “break point” means “if the receiver wins this point, they win the game.” Break points are notable because the default assumption in tennis is that it’s easier to win a game when you are serving than when you are receiving. So if you win a game as a receiver (or, in tennis jargon, if you break your opponent’s serve), it’s generally viewed as an indication that you are pulling ahead (or making a comeback).

I think TJR understands this, but there are also some uses of break point that I cannot interpret. Here’s another excerpt.

“Serve the ball, Carrie.”
With lightning speed, I toss it up in the air and slam it across the court. It flies from my racket straight across the net and into the ground. It bounces out of reach before Bowe gets to it.
“Break point,” I say.
“Goddammit,” Bowe says as he throws his racket.
I shake my head and start walking to the bench to drink some water.

This isn’t a break point! First, Carrie is serving, so if she gets into a position to win the game, it’s not a break point. It’s just a game point. Second, after the point, they both go back to the benches to drink some water, which suggests that Carrie’s ace has actually won her the game. I truly cannot understand what Carrie saying, “Break point,” is supposed to communicate here.

There are some other stray jargon issues I noted: using return for shots other than the return of serve, using groundstroke to indicate a kind of shot (nobody would say, “hits a groundstroke,” they would specify whether it was a forehand or backhand), and referring to winning “in two sets” when the only way people actually say that is “straight sets.” To me, these are facts about tennis and how tennis is discussed that are not up for interpretation or change, even in a fictionalized world.

Heading West from LA to Indian Wells…

OK there is just one other thing that really made me scratch my head. Carrie goes to the Indian Wells tennis tournament as a spectator to do some scouting. Indian Wells is one tier below Grand Slam— it’s a huge tournament where you can expect every top 10 player in attendance. It takes place in the California desert (the town is called Indian Wells, and it’s near Palm Springs) and from LA, the only way to get to Indian Wells is to drive EAST. When Carrie, her agent Gwen, and her dad leave LA to go to Indian Wells, Jenkins Reid writes: “[We] pack our suitcases into Gwen’s SUV and head west for Indian Wells.” I don’t think it’s reasonable to assume that Jenkins Reid was simply imagining a fictional version of California where Indian Wells is west of LA— I think she just doesn’t know!

You can’t make this up!

A straw man view of fiction would be to say that it’s “made up” or “not true or real.” Certainly some part of every work of fiction IS real (e.g., the characters respirate, have organs, talk to each other, …), and furthermore, there are entire genres of literature that focus on retelling or embellishing historical events. Fiction that is set in the actual world (or something very close to it) borrows parts of the real world to help us know the world of the book. So why does it seem to be the case that some parts of the real world (like the scoring system of tennis or the location of Indian Wells) must be kept?

The best I can come up with is that there needs to be a reason to make something up. Fictional characters are perhaps the easiest to accept because we know that we’re reading the novel to connect with characters (and probably also because most books have invented characters!). Historically inaccurate statistics may be too much for some to accept (as some reviews of this book contended), but we can also accept them if we understand/believe that Carrie Soto is remarkable— maybe this is a step away from magical realism (ha)? But the inaccurate scorelines and use of jargon in Carrie Soto serve no discernible purpose. So, while it should be possible for me as the reader to say, “Oh, the world of this book is different because it contains players like Carrie Soto and Bowe Huntley and also because tennis is scored differently and Indian Wells is in the Pacific Ocean,” that’s not where I go. I just think Taylor Jenkins Reid and her editor(s) don’t know enough about tennis and didn’t care about getting those details right (which, fine!).

Add this to your list of stories to lean on when you start to hear your inner perfectionist

Taylor Jenkins Reid is certainly a successful novelist, and this book was reasonably popular (though not popular enough to have its own Wikipedia page). I finally bought myself a copy after seeing that the ebooks at my public library were constantly checked out— people are reading this! And even as a tennis fan, I would still say I liked it enough! And look at all these thoughts I had about it!

Even though the book isn’t really “about tennis” in the end, I still think it’s interesting that these mistakes made it into the final story. That TJR was writing about Indian Wells and wrote “head west” and then didn’t say, “Oh, I better check whether that makes sense.” So the lesson for perfectionists out there is: it’s fine! You can write a successful book involving tennis even if you don’t really know that much about the game. Some people will care that you got some details wrong, but also plenty of people won’t care at all. So just remember that!

And, lastly, it made me curious about reading some tennis-based fiction where the tennis knowledge is actually correct! I’ll have to look into that…

How I decided to leave my academic job

The shortest and simplest explanation is that I left my academic job at OU to prioritize my personal life, and this is how I described it at first. I was in a long-distance relationship for the entirety of my time at OU, and my partner and I were ready to share a home base. But that explanation is too brief— it took years for me to reach this decision, and I didn’t make it lightly. Inspired by some friends who have recently shared more details about their departures, I decided to share a longer version of my story.

I want to acknowledge up top that I had a decent savings by the time I left (Thank you, Oklahoma!), another financial safety net for the job transition (my partner), and no family to take care of or other big responsibilities (Cookie is sooo tiny!). All of this released a lot of worry and uncertainty around my decision. I’m grateful for that, and I wish that everybody could choose what’s right for them without having to worry about how to make ends meet.

Settling into the field in graduate school

To be honest, there were indicators during my academic career that I wasn’t all-in. When I decided to attend UCSC, I specifically opted for on-campus housing so that if I decided to leave the program, there would be no penalties for breaking my lease. In my first year, I felt insecure about my level of dedication. I can’t really access what I was feeling at the time (it was, of course, FIFTEEN DANG YEARS DANG AGO), but it was something to the effect of “My life does not revolve around this.” I liked linguistics, but I liked other things, too! A new friend, who was in the MA program at UCSC and was planning to leave the field, used the phrase “I want to be a body, not a brain!” to describe their feelings, and I thought, “Me too! I’m also a body in this world!” I brought this up at a social event among grad students in a “Don’t we all feel this way?” kind of way, and I was surprised by the response. Apparently, we did not all feel this way. But I stayed in the program (I like linguistics!) and continued to try my hardest to do well.

By the time I started to work on concord (in my third year), I had really settled into things. It was exciting to feel like I was breaking new ground, and even though I think I probably cried about once a week (lol), I recall enjoying my life and enjoying my work. When it came time to apply for jobs, it did not even cross my mind to go into industry, even though there were many recent graduates from my program who had found gainful employment in the tech industry. I was certain that I had to try to be a professor even as I was uncertain if it was what I wanted. In fact, I didn’t really consider asking my partner to move with me to Oklahoma, and one of the reasons was that I wasn’t sure if I wanted to be a professor for the rest of my life.

Add “decide your whole life” to my pre-tenure expectations

When I was getting ready to head to OU, my partner’s career was also starting to go in new directions, and we mutually decided to let each other follow these career paths without the pressure of making a decision. I spoke with faculty at OU and at other departments who would console me by saying things like, “Oh, such-and-such faculty member lives far away from their partner, and they make it work!” and I thought, “But this is not something you HAVE to deal with in life! And it is not what I want!” I knew, at least, that sacrificing sharing a home base with my partner was not something I was willing to do forever.

Maybe because of that, I put a lot of pressure on myself to figure out if being a professor was what I wanted in those early years. Asking my partner to give up his career to move to Oklahoma (or somewhere else, if I got a job elsewhere) was a big deal to me, and I wanted to be sure. But of course, I had never been sure (you know, like sure-sure). And furthermore, my early years at OU were difficult because I was having trouble getting work through prepublication peer review (which is a bad practice). It felt impossible to decide if this was truly what I wanted when I was also struggling to feel successful. 

Over winter break in my second year, I had a panic attack while visiting my partner, precipitated by stress/anxiety about everything I needed to do for work. This was an important signal, as I have had maybe three panic attacks my entire life. I decided I should start seeing a therapist again.

One of the early revelations that came out of therapy was that I was (i) not fully committed to my position at OU (because I was “trying to figure out” if it was what I wanted) but (ii) acting in my role as though I was fully committed in terms of effort expended and expectations I had for myself. This made it very difficult to feel calm and secure, so I decided that I had to commit to my faculty position and my life in Norman as long as I was there. I bought a house at the end of my second year (a house I loved dearly; Thank you, Oklahoma!), and I released myself from the pressure of needing to figure it out. The next several semesters were generally pretty fun, I had some more success, and I liked my job.

“Wondering whether” can be an answer

In roughly the middle of my fourth year, I was still wondering if it was what I wanted. I realized that if I was still wondering after 3.5 years, that was enough of an answer. I also did the math and realized that I had 30+ working years left (Thank you, American capitalism!), which is plenty of time to build a second career. I liked my job enough, but I didn’t feel like it was my calling or passion, and I knew what I was sacrificing in order to pursue it. I began to think that it was worth it for me to try to find something else that either (i) I liked even more or (ii) did not ask me (or my loved ones!) to sacrifice so much. 

After a particularly stimulating academic conference, I reconsidered the decision for the final time: Did I want to leave Academia entirely, or did I just need to leave my job at OU for a “better” academic job? Another friend who had left for industry reminded me in a helpful phone conversation that the academic job that I could be happy with over the long term was extremely difficult to get (if it even existed). For years, I had held on to the promise of an academic job where (i) I was treated well by colleagues and the institution (and that includes pay and research funding), (ii) I could research and teach what I wanted, (iii) I had dedicated and funded graduate students, and (iv) I lived in a place I wanted to live. But how many of those jobs are there? So many faculty members sacrifice some number of these things in order to be faculty. I had been at an institution where I was sacrificing to some degree on all four of those qualities, and though there were things about my work that I liked, I knew that they were not enough to endure those sacrifices indefinitely.

It would take time to find a new way to make money, but I had over 30 years to build a new career. I resolved to go for it.

What do I think about it now?

While my relationship wasn’t the reason I left academia, it was the catalyst for my decision. It forced me to wrestle with issues much quicker than I might have otherwise, but I ultimately arrived at the right decision for me. Even if another academic job existed in San Francisco, where my partner and I planned to live together, I knew that I wouldn’t apply for it. (And I still wouldn’t.)

Folks sometimes ask if I miss it, and the answer is yes and no (but mostly no)! I say this every time I talk about it, but I’m so glad that I spent 11 years of my working life getting paid to study generative linguistic theory and teach linguistics. Human language is one of the great loves of my life, and I find it incredibly fascinating and fun. I do not regret my decision to take the job at OU, I don’t regret my decision to stay for five years, and I don’t regret my decision to leave. I have said before that I don’t miss the amount of research I used to do, but I also don’t really miss teaching. That doesn’t mean I didn’t like it! It just means that there are so many things in life to be enjoyed, I found other things to enjoy, and it was enough to fill my cup. I treasure the good relationships with students and faculty I made while I was at OU, but I definitely do not miss grading papers. So it’s “yes!” because I remember it with fondness, but “No!” because I don’t regret leaving, and I don’t yearn for the work.

I was so worried about what the rest of my life would look like if it didn’t revolve around linguistics, teaching, and the broader linguistics and higher education community, but it’s okay. I still have friends from that time, and I still engage with linguistics research from time to time. But now, it’s entirely on my terms.

(This is the third post in a series of blog posts about my transition from academia to industry and my feelings about my time as an academic teacher and researcher. To read more, see the “industry transition” tag.)

Reflections on research after academic jobs

When I left the job where I was paid to do linguistic research (among other things), I told myself that I was going to continue to do research (i) as long as I had time for it and (ii) as long as I found it fun (or you know, as long as I wanted to). I anticipated that within 2-5 years, the sun would set on my ability to contribute new research to generative and typological linguistics. I left my academic job just about 2.5 years ago (in May 2019), and I have been working in industry for almost 11 months. As I continue to develop roots in this next phase of my working life, these thoughts have been creeping back into my mind.

Between deciding to leave and actually leaving: anticipatory grieving

Some post-academic folks I speak to seem to have little love lost over their academic research interests, but this was not me. I felt very sad about saying goodbye to those things. In particular, I recall feeling saddest about the changing relationship to Estonian, my primary research language. Leaving my academic post meant saying goodbye to biannual trips to Estonia (for fieldwork and swimming), and it meant less occasion to contact the friends who taught me about their language. I might say that I was worried I would miss Estonian and Estonia so much that I would regret my choice to leave my academic job, but really I think I was just sad that things were going to change.

In the middle: sometimes a source of comfort and purpose and sometimes a distraction

My transition to industry took longer than I expected: 14 months after I relocated to SF (11-12 of those spent actually searching), I got my first industry offer. I have shared this before, but it bears repeating (just so people know): I experienced some of the lowest/darkest moments of my life during that time. Searching for jobs ALWAYS sucks, and so does trying to make a career transition. During this period, doing a little bit of linguistic research would help temper the feeling of purposelessness I often felt. For example, for a few months, I had a weekly reading group with Ruth Kramer, who has been both a dear friend and research advisor essentially since we met in 2008. I spent maybe 20-30% of my “working time” doing linguistics, just because it gave me something to do that I felt like I knew how to do.

There were other times where doing linguistics felt like an indulgence. It felt like it wasn’t quite the thing I was “supposed to do” in order to make myself more competitive for a job. In retrospect, it wasn’t that doing linguistics was EITHER helpful for me or a distraction; it’s that sometimes I needed it, and sometimes I didn’t.

Now completely in industry: I liked it then and I like it still, but I don’t regret my choice

I’m now just over 6 months into my first permanent position at industry, working on problems I find interesting with a team of people I truly enjoy. I did actually have some research output over the last 12 months, and (to my surprise, honestly) I was recently invited to give a colloquium and contribute to another handbook, so I think it’s clear that I’m still doing research at this point. BUT GOODNESS, it’s even harder to make time for it now! After I wrapped up my joint paper with Kyle Mahowald and Dan Jurafsky, I didn’t have any research deadlines, and weeks without doing linguistics passed by before I realized. It’s not that I no longer enjoy it, it’s just that it’s one of many things I enjoy that I have to use time outside of work to enjoy.

This made me think about how I felt before I left: would I miss it? Would I be sad about letting these things go? At this point in my post-academic life, it seems the answer is “No.” That could be because I’ve let go gradually. It could also be because I still haven’t completely let go— I have a handbook chapter that is still set to come out (handbooks are… slow) and I was just asked to contribute to a different handbook (hello, deadline). But I think a large part of it is (i) I’ve had space to move on and (ii) I have a new career that is providing plenty of intellectual stimulation. AGAIN, I must stress that this doesn’t mean I didn’t like it then or don’t like it anymore! I’m still very happy I spent 11 years of my working life dedicated to linguistics teaching and research. I’m also happy about learning to do new language-related things!

Future: What’s actually worth my investment? Can I walk away from unanswered questions?

Since leaving, I have realized that even if I continue to do theoretical/typological research when it is no longer part of my job description, it will not look the same as it did before. There were many research-related activities I did when I was a professor:

  • Read theoretical papers: both to stay current and to try to find inspiration when solving a particular puzzle
  • Write papers: to share knowledge and proposals in a permanent form
  • Present at conferences: to share knowledge and proposals
  • Give invited talks: both colloquia and working group talks
  • Review articles: if I’m going to keep writing, I should keep reviewing

This is a lot of tasks! And realistically, on a busy research week, I probably can spend about 5 hours on this. Deadlines have become significantly more motivating than they were in the past. I have been able to complete necessary work and not much else. For example, I have had to be more selective about reviewing, and I barely read enough to support my own projects. Forget about staying current!

At some point, it will be time to effectively stop. I have a pipe dream of writing a book and just making it available, published or not. There are too many things I’ve learned—especially about concord—to just leave them in my brain. There are questions that I want to know the answers to, and if I don’t find the answers to these questions, then I don’t get to know what they are (because either nobody else will, or they they won’t tell me if they do). Trying to get all of my knowledge on paper is one way to possibly avoid that, but I also think I will have to leave some of these questions unanswered. I suppose that’s also just part of moving on from jobs more generally— letting go of in-progress work.

What I (a linguist) did while searching for a job (in tech)

This is a post about (or the first post in a series about) the search for my first job in the language + tech industry space. I perhaps could have called this something like What I Did to Get a Job, but one of my opinions about getting your first industry job is there is no silver bullet for the task, so I can’t say that any of these things were necessary in getting my first job. However, I did some things while searching for a job and then did manage to get a job (first a contract job and then a full-time job).

To say that again another way:  there is no one thing that you must do to be able to get a job nor is there one thing that will enable you to easily find a job if you do that one thing. Unfortunately! So what you should do is just keep trying to develop.

Coding

What I did to learn/practice coding (mostly Python):

  • Free intro course (keyword FREE); I did Codecademy‘s Python 2 course, because their Python 3 course costs money.
    • There are differences between Python v2 and v3 but they are minor and you will pick them up.
    • This will help you learn the basic data structures.
  • Automate the boring stuff is a good book, and all the chapters are available online (just scroll down the linked page):
    • Learn to manipulate files (reading/writing CSVs, text files, and JSONs): Chapters 9, 13, 14, 16 (at least!)
  • Free mini-courses on Kaggle
    • They have great bite-sized courses on a variety of topics, e.g., Intro to Python, Pandas, intro to ML, advanced ML, data cleaning
    • When you’re done, you can even add a certificate to your LinkedIn profile, which you should do!
  • NLP-focused content
    • Introduction to SpaCy (great library for NLP)
    • NLTK book
    • My two cents: don’t worry about the syntactic details; focus on internalizing the steps of an NLP pipeline in a broad sense
      • I got overly concerned with knowing this well—I never really got there, and I also haven’t had to know it well for either of my jobs or for my interviews.
      • The file management stuff is more important!

The main thing I have done at my jobs with Python is file and data management. For file management, learn to do things like (i) list all the files in a directory, (ii) perform the same operation on every file in a directory (or just every .csv or .txt file…), etc. For data management, you want to be able to manipulate different file types; I would say focus on being able to import and export csvs and jsons. In my opinion, you should prioritize learning (some of) the pandas library, because it’s powerful and also used by data scientists. I have some aspects of pandas memorized by now, but I still look things up all the time. That being said, I once had a coding interview where the code was written with the csv package, and I couldn’t remember the syntax. I just tried to do my best in the interview, but I did feel afterwards like I should make sure to remember how the csv package works. For language work, the csv package is also probably quite a workable solution— some of my colleagues at both jobs have used csv instead of pandas.

Learning ML/AI/NLP

The key thing here is that you need to focus on being conversant in these concepts rather than necessarily being able to write them. Maybe you can get to that, too, but step one is understanding what the pieces are. The reason to watch and read these things is not so you can necessarily do this work (speaking for myself, I do not build ML models at work), but so you know enough about them to know kind of how the enterprise works. To be clear, there are some “linguist” jobs where you build models, so if you’re interested in it, it’s definitely worthwhile. However, you can secure a job without building a model. Again, these are the things that I read or watched while I was searching, not necessarily “must watch” or “must read” sources!

We did not get to see the models in my job at Amazon. I still think it was helpful for me to know a little bit how these programs work so I could talk/think about how the data would be used. 

“Do a project”

A lot of people told me some version of “do a project.” If your feeling when you hear that is “Sounds good, but what?” I don’t blame you! As someone offering advice, it is easy to think that project ideas abound, because once you start working, so many more ideas come to you. However, when I was searching, I just didn’t know what projects I could do. Most identifiable NLP projects (i.e., papers) are multi-authored, so I felt that I couldn’t make headway there (let alone the fact that I was still learning!). I definitely didn’t think I was doing any projects that were big enough to count as projects (whatever that means), and at times, I felt like I couldn’t even if I wanted to.

Surprise! I actually did some projects

Now that I’ve been working about 8 months, I can see that I did “do a project” a few times. Here is a reasonably complete list.

  • Poketext: scraped Pokédex entries from Bulbapedia pages and saved it in one giant text file
    • When I started this, I had in mind that I was building a corpus. I don’t know what the corpus would be used for— probably a good idea to have an answer to that.
    • This gave me some experience web scraping with the Python module called BeautifulSoup
  • Poketext (part deux): try to fix issues in some of the sentences that I saw using SpaCy
    • Some of the sentences in the Pokédex would refer to the Pokémon by name, but some would just use a pronoun or an NP like “this Pokémon”.
    • I created an algorithm to replace the pronouns/NPs with the name of the Pokémon, and I encountered some interesting issues in the process.
  • Concord: I converted my typological database from a somewhat unwieldy spreadsheet to a format that was more computationally friendly and easily updated.
    • Each individual language in the data set is one JSON file
    • I wrote scripts to import all the JSON files from the directory and output descriptive stats about the data set
    • I wrote a program that creates a data file for any language that is new to the data set (after I’ve documented it myself). It asks me a few questions about the language and then creates the JSON file.
    • Because I did this, Kyle Mahowald found my data and we started a research collaboration.
    • FWIW: this is the project that ended up being the most useful (imo) for getting my first job.
  • Estonian spell checker: Messed with some Estonian data to adapt Peter Norvig’s very simple English spell checker for Estonian.
  • My blog!: Though I do sometimes write about esoteric issues connected to my linguistic research, I also tried to write some posts addressing language + data in a friendly, accessible way.

My recommendation is that you put these projects on a github and/or on a website, no matter how small you think they are. If recruiters or hiring managers get curious enough about you to look for your web presence, you want to give them something to chew on. Projects that are representative of your current skills make the most sense—if they’re not flashy enough for a particular job, then trust me: you don’t want that job (right now)!

Additional project ideas

Feel free to riff on any of the project ideas stated above. If none of those sound fun to you, here are some ideas to inspire any do-a-projects you might pursue.

  • Build a test set for an imaginary classifier: Test sets are smaller than training sets, so it won’t take you as long to annotate them.
    • Columns (one idea): label/annotation, URL, first paragraph, first paragraph tokenized or split on spaces (eg, using SPLIT() in Sheets)
    • Classifier ideas:
      • X or not X: news articles that are or are not about a certain thing (eg, about accidents or not, about the stock market or not, or even subjective categories)
      • native vs non-native (or: native vs. translated): sentences/paragraphs that are or are not from native speakers.
  • Build a corpus with web data (possibly scraping with BeautifulSoup if you can’t find the data more easily): collect examples with the idea that a person could design an annotation using the data you’ve collected.
    • There are lots of “corpus of movie review” tutorials online— these won’t “make you stand out”, but you will learn a lot from doing them.
    • I also saw something that involved pulling all the dialogue from a TV show and using that as data
    • Could use this corpus to feed into a project like the one above.
  • Playing with old research data: If you have any research data that can be put into spreadsheet format, try to do something with it in python.
    • Maybe that’s data visualization,
    • maybe that’s creating randomized samples of the data,
    • maybe that’s writing a script that would let you add to the data (e.g., what fields does the script need to ask for),
    • maybe you want to create a database of linguistic examples that you have gathered in your fieldwork/research and tag it for relevant information (eg, these are my relative clause examples, these are my wh-question examples, …)
    • If you don’t have any spreadsheet data from your research, you could download some samples from wals.info and play with those. Eg, can you figure out how to download 3 samples from WALS and combine them all into one file (so that if a language appears multiple times, you consolidate all of its values into one row?). This stuff isn’t super hard to learn!
    • just do something to give yourself something to work on and slowly figure out

The important thing to remember: there’s no secret beyond trying

Unfortunately, if you don’t already have computational or statistical skills, it can be hard to show a recruiter that you can contribute. There is no secret to this task— you just have to maintain your ability to try new things and keep waiting for a little bit of luck. I don’t mean this in a hokey All You Need to Do Is Try kind of way. I just mean it’s probably not helpful to worry about whether you have done The One Thing You’re Supposed to Do. Such a thing probably doesn’t exist. When I wasn’t worried about that possibility, I kept working on projects as long as I found them fun/interesting/whatever, and when I stopped feeling excited about them, I moved on to something else. I never did finish that ML course even though I told myself I would. It’s okay— just keep trying to practice or learn.

A perspective on data usability from the concord typology project

Cross-linguistic studies involve a lot of data collection. If we want to make the most progress towards understanding something, it makes sense that we should make our work usable for people besides ourselves. We might be able to get some additional help!

When I started the concord typology project as a researcher, I knew I wanted to keep track of my data in such a way that other researchers could easily build on the work I had done. I also wanted to make sure that anybody who wanted to retrace my steps would be able to without having to build the study from scratch.  After writing the proceedings paper based on the initial results, I uploaded the data (with the paper) to OU/OSU/OCU’s SHAREOK archive. Here’s what’s in there:

  • Main article: the proceedings paper (geared towards academic audiences)
  • Research data (spreadsheet): all the coding and classification I did based on the data that I collected; easy to digest reasonably quickly
  • Research data (read me): an explanation of the contents of the archive
  • Research data (examples + examples appendix): the actual linguistics examples that the spreadsheet is based on; sort of like research notes and thus not as easy to digest

I spent some time cleaning up the data to prepare it for eyes other than my own, but it’s hard to perfect it on the first pass. I called it good enough and then got back to work collecting more data.

This is an attempt to use a slightly blurry and slightly purple clipping from the research data spreadsheet as an artistic way to break up the flow of text. Hashtag data is art 😆

One year later, someone finds and uses the data!

About a year after I published my data in the archive, I got an email from Kyle Mahowald, an assistant professor of linguistics (at UC Santa Barbara) who is interested in computational modeling of cross-linguistic studies (like mine). He stumbled upon my data and decided to start building a model based on the data that could account for issues of genetic and geographic proximity. This was, of course, very exciting: somebody was building on the work that I started! I spent a good deal of time getting the data ready for the archive, so seeing that somebody found it and used it made all that effort worthwhile.

This brings me to my next point: when you think about making your data usable, consider the user carefully, and make sure the data is as usable as possible. In the version in the SHAREOK archive, I made a choice that negatively affected usability. When Kyle initially wrote the model, it was making a few predictions that were strikingly different from my results. The issue arose from the coding schema I used for for the spreadsheets. In brief: there was an overlap in some of the labels I had used, and so the script was treating some distinct labels as though they were the same. The bug was an easy fix, but we only noticed it because Kyle and I started collaborating and discussing the model he developed. It got me thinking about how I could make my data not only available, but (even more) useable.

Hammer, Sledgehammer, Mallet, Tool, Striking, Hitting
If you give a very simple program a hammer, it’s gonna start looking for nails (or whatever this smashed thing is).

Iterating towards better usability

After my initial conversations with Kyle, I put some work in to improve the usability of my data. I wanted something that achieved a balance between these three things:

  1. Easy to read (for humans)
  2. Easy to process (for computers)
  3. Easy to update (for me)

For this, I settled on storing the coding for each language as a JSON file (as I mentioned in this post). I find JSON files relatively easy to read—especially if you save them with some formatting for readability—and they can be easily converted to other formats. I wrote a few scripts to convert the existing overlapping labels into a system without overlap. And to add new data, all I have to do is add a JSON file for the language I just documented (which I’ve already written a script for).

I now store and update the data on OSF, a free and open platform for sharing research. This means anybody (even you, dear reader) can download the current state of the study this very moment! If you don’t have experience working with JSON files yourself, don’t worry: I have a script on my github that processes the important data from all the JSON files and saves it as a single CSV. So, even if you haven’t used anything besides Excel/Sheets, you can still look at the data!

Keep your data user-friendly!

To sum up, the big lesson here is to keep your data user-friendly. If you want whoever uses your data—be they a research colleague, a coworker, or a client—to be able to build on the work you’ve put in, think about how they might use the data and try to make that as easy as possible.

Books, Pages, Story, Stories, Notes, Reminder, Remember

A very simple spelling corrector for Estonian

If you’ve spent any time looking at online NLP resources, you’ve probably run into spelling correctors. Writing a simple but reasonably accurate and powerful spelling corrector can be done with very few lines of code. I found this sample program by Peter Norvig (first written in 2006) that does it in about 30 lines. As an exercise, I decided to port it over to Estonian. If you want to do something similar, here’s what you’ll need to do.

First: You need some text!

Norvig’s program begins by processing a text file—specifically, it extracts tokens based on a very simple regular expression.

import re
from collections import Counter

def words(text): return re.findall(r'\w+', text.lower())

WORDS = Counter(words(open('big.txt').read()))

The program builds its dictionary of known “words” by parsing a text file—big.txt—and counting all the “words” it finds in the text file, where “word” for the program means any continuous string of one or more letters, digits, and the underscore _ (r'\w+'). The idea is that the program can provide spelling corrections if it is exposed to a large number of correct spellings of a variety of words. Norvig’s ran his original program on just over 1 million words, which resulted in a dictionary of about 30,000 unique words.

To build your own text file, the easiest route is to use existing corpora, if available. For Estonian, there are many freely available corpora. In fact, Sven Laur and colleagues built clear workflows for downloading and processing these corpora in Python (estnltk). I decided to use the Estonian Reference Corpus. I excluded the chatrooms part of the corpus (because it was full of spelling errors), but I still ended up with just north of 3.5 million unique words in a corpus of over 200 million total words.

Measuring string similarity through edit distance

Norvig takes care to explain how the program works both mechanically (i.e., the code) and theoretically (i.e., probability theory). I want to highlight one piece of that: edit distance. Edit distance is a means to measure similarity between two strings based on how many changes (e.g., deletions, additions, transpositions, …) must be made to string1 in order to yield string2.

A diagram showing single edits made to the string <paer>
Four different changes made to ‘paer’ to create known words.

The spelling corrector utilizes edit distance to find suitable corrections in the following way. Given a test string, …

  1. If the string matches a word the program knows, then the string is a correctly spelled word.
  2. If there are no exact matches, generate all strings that are one change away from the test string.
    • If any of them are words the program knows, choose the one with the greatest frequency in the overall corpus.
  3. If there are no exact matches or matches at an edit distance of 1, check all strings that are two changes away from the test string.
    • If any of them are words the program knows, choose the one with the greatest frequency in the overall corpus.
  4. If there are still no matches, return the test string—there is nothing similar in the corpus, so the program can’t figure it out.

The point in the program that generates all the strings that are one change away is given below. This is the next place where you’ll need to edit the code to adapt it for another language!

def edits1(word):
#     "All edits that are one edit away from `word`."
    letters    = 'abcdefghijklmnopqrstuvwxyz'
    splits     = [(word[:i], word[i:])    for i in range(len(word) + 1)]
    deletes    = [L + R[1:]               for L, R in splits if R]
    transposes = [L + R[1] + R[0] + R[2:] for L, R in splits if len(R)>1]
    replaces   = [L + c + R[1:]           for L, R in splits if R for c in letters]
    inserts    = [L + c + R               for L, R in splits for c in letters]
    return set(deletes + transposes + replaces + inserts)

Without getting into the technical details of the implementation, the code takes an input string and returns a set containing all strings that differ from the input in only one way: with a deletion, transposition, replacement, or insertion. So, if our input was ‘paer’, edits1 would return a set including (among other thing) par, paper, pare, and pier.

The code I’ve represented above will need to be edited to be used with many non-English languages. Can you see why? The program relies on a list of letters in order to create replaces and inserts. Of course, Estonian does not have the same alphabet as English! So for Estonian, you have to change the line that sets the value for letters to match the Estonian alphabet (adding ä, ö, õ, ü, š, ž; subtracting c, q, w, x, y):

    letters    = 'aäbdefghijklmnoöõprsštuüvzž'

Once you make that change, it should be up and running! Before wrapping up this post, I want to discuss one key difference between English and Estonian that can lead to some different results.

A difference between English and Estonian: morphology!

In Norvig’s original implementation for English, a corpus of 1,115,504 words yielded 32,192 unique words. I chopped my corpus down to the same length, and I found a much larger number of unique words: 170,420! What’s going on here? Does Estonian just have a much richer vocabulary than English? I’d say that’s unlikely; rather, this has to do with what the program treats as a word. As far as the program is concerned, be, am, is, are, were, was, being, been are all different words, because they’re different sequences of characters. When the program counts unique words, it will count each form of be as a unique word. There is a long-standing joke in linguistics that we can’t define what a word is, but many speakers have the intuition is and am are not “different words”: they’re different forms of the same word.

The problem is compounded in Estonian, which has very rich morphology. The verb be in English has 8 different forms, which is high for English. Most verbs in English have just 4 or 5. In Estonian, most verbs have over 30 forms. In fact, it’s similar for nouns, which all have 12-14 “unique” forms (times two if they can be pluralized). Because this simple spelling corrector defines word as roughly “a unique string of letters with spaces on either side”, it will treat all forms of olema ‘be’ as different words.

Why might this matter? Well, this program uses probability to recommend the most likely correction for any misspelled words: choose the word (i) with the fewest changes that (ii) is most common in the corpus. Because of how the program defines “word”, the resulting probabilities are not about words on a higher level, they’re about strings, e.g., How frequent is the string ‘is’ in the corpus? As a result, it’s possible that a misspelling of a common word could get beaten by a less common word (if, for example, it’s a particularly rare form of the common word). This problem could be avoided by calculating probabilities on a version of the corpus that has been stemmed, but in truth, the real answer is probably to just build a more sophisticated spelling corrector!

Spelling correction: mostly an English problem anyway

Ultimately, designing spelling correction systems based on English might lead them to have an English bias, i.e., to not necessarily work as effectively on other languages. But that’s probably fine, because spelling is primarily an English problem anyway. When something is this easy to put together, you may want to do it just for fun, and you’ll get to practice some things—in this case, building a data set—along the way.

How to build a cross-linguistic database

One of the best things about being a linguist is that language data is all around us. We interact with it in so many ways every day: listening to podcasts, talking to housemates, and even shopping at grocery stores. It’s easy to start paying closer attention to the language we hear and see, but how can we learn from all this data to better understand how language works? In this post, I’m going to focus on one aspect of this: understanding how other languages work by building a cross-linguistic database.

To do this, I’ll talk about a specific example, a certain kind of agreement with nouns (which I call nominal concord, but I’ll just call it agreement with nouns in this post). The English words this/that change their form based on whether the noun they modify is singular or plural: e.g., these bananas vs. this banana. Other languages do similar things. Many people who have studied a Romance language like Spanish remember that articles and adjectives (among other things) have to match the gender of the noun they modify: e.g., la casa blanca ‘the white house’ vs. el edificio blanco ‘the white building’. How do we figure out the properties of this process so that we can use it to understand language better?

A map showing the presence of concord (i.e., agreement with nouns) in the world’s languages. Data from Norris (2019), map created in R with the lingtypology package.

Collecting language examples for the database

Since you’re already swimming in data, you could approach data collection in a very grassroots way. You could travel, taking pictures of language you see in the world, or you could take screenshots of other languages you encounter on blogs or social media. But this can take a long time when you’re looking for data that shows you a specific thing. To develop the database more rigorously, we can turn to more formal sources. We can use grammars, which are book-length reference guides for the properties of languages. Some of these are freely available as books (e.g., Language Science Press, Pacific Linguistics) or PhD dissertations (especially low-resourced languages). Some language data is already compiled and available in online databases (like World Atlas of Language Structures, or Universal Dependencies Corpus) or established corpora (like the Corpus of Contemporary American English, or for something completely different, Estonian corpora available at keelveeb.ee).

There is an important aside here: if we want to understand how language in general works, we have to make sure we’re not looking too closely at one particular language or language family. The languages that most people in North America and Western Europe are familiar with are Indo-European languages. Because these languages are so familiar, there is a reflexive tendency to view their properties as normal or common. To go back to the example of agreement with nouns, we might think that it’s normal for articles and adjectives to agree with noun, because that’s what they do in Spanish. It’s important to remember that we don’t actually know if that’s true! The only way to know for sure is to build a database that does not have overrepresentation of one language or language family.

Selecting a feature/tag set

Now that we’ve talked about sources you can use for your data, the next thing to determine is the set of features or tags that you’ll use to structure the data you collect. Naturally, you will want to focus on features that seem relevant for understanding whatever you’re looking at. If you were looking at agreement with nouns, you could catalog which words agree with nouns and what properties they agree with.

So, after collecting your data and storing it (e.g., in a Google Doc), you would also record what words agree with nouns in that language and what properties were relevant for the agreement. There are many features of a language that are likely irrelevant for what you’re looking at. You might need to do some exploratory analysis first to get the lay of the land before you decide on your feature set.

Managing the data and features

I mentioned that you might choose to store the examples you collect in a Google Doc. What about the features? There are several options for managing databases of this type. A non-exhaustive list of options:

  1. Spreadsheet (e.g., Excel/Sheets): easy to read; allows sorting “by hand”
  2. CSV: a bit harder to read, but can be easily opened by a spreadsheet program or fed into R or Python for more sophisticated computational or statistical analysis.
  3. JSON: easier to read with human eyes (in my opinion), easily read with all sorts of programming languages

Saving the data in a format that can be easily read by something like Python is a good investment. You can write scripts to read in all the data you’ve collected and then tabulate any numbers you find relevant. The approach I use is to save each language as an individual JSON file so that updating the database is simple: all I have to do is add the new JSON file(s) to the proper directory. Then I can run the scripts I’ve written to see how the database has changed.

Image showing a sample of a JSON file
A piece of a JSON file containing information about agreement with nouns in Finnish.

I know I went through this part pretty quickly—I’ll share more about it in a subsequent post!

Once we have enough languages in the database, we can start to extract insights about how different languages do or don’t behave similarly, and through that, we start to learn about language in general. When I built a database like the example I have been discussing in this post, I learned that demonstratives and adjectives are the most likely categories to agree with nouns. In fact, if demonstratives agree with nouns in a language, then it is likely that adjectives will, too. This is just one way in which English turns out to be weird: demonstratives (this/that) agree in English, but adjectives don’t—we don’t say heavies books!

Congratulating yourself

That’s that! As with lots of language work, the most time-intensive part is not really analyzing the results, it’s ensuring that the analysis is based on good data. Collecting good data can take a long time, especially when you’re pulling it from a lot of different languages. So, if you feel compelled to make a database like this one, you might as well start now! A database that would be very useful for both academic and applied contexts would be one that catalogs different word orders that can be used based on the discourse context (e.g., topicalization or focusing) —there are related databases on WALS but they’re less pointed. Any linguists reading this will know that that database could be quite difficult to construct!

A simpler task would be to contribute to an existing database. The Universal Dependencies Corpus is a great example. Right now, the most developed samples in the database are Indo-European languages (and a few other major world languages of Asia). As a result, the Universal Dependencies Corpus is unfortunately still biased!

Some language in the wild from Estonia— the translation of Thoreau’s “Why should we live with such hurry and waste of life?”
css.php