Tuesday, March 29, 2011

Announcing Rogue Editing & Design

Hey, folks; Susan here, to let you know that I've launched a new site for my knitting and generally yarn-related activities! I may still post non-yarn-related things here on Two Piece Set, although as you've noticed traffic is pretty low around here these days, so I don't expect you to hold your breath. Instead, follow me over at www.RogueEdits.com, especially if you're looking for a tech editor or some fun new knitting patterns!

Here's what the new site offers:

Note that I still don't accept Google-related questions on my personal sites, either here or at the new site. For anything Google-related please contact me and/or my co-workers through our Help Forum.

Tuesday, January 25, 2011

New Year's resolutions & OKRs

Hello, 2011!

Every quarter at Google we write these things called OKRs: Objectives & Key Results. Each individual has OKRs, each team has team-level OKRs, and Google itself has company-wide OKRs. Objectives are big-picture things that you want to accomplish: achieving world peace, making users happy, becoming an awesome roller derby skater. Key results are measurable ways that you will quantify your progress toward said objectives. Key results for achieving world peace might include reducing the number of people killed weekly in conflict zones by x%, or reducing nuclear weapons stockpiles by y%. Key results for becoming a rollergirl might include getting onto a roller derby league, or getting your 25-lap time down to ≤ 4 minutes.

There are several things I like about the OKR system. It not only requires you to set goals for the quarter, but to think about why you're picking those goals. What objectives are you trying to achieve? Just making a to-do list of things you need to finish this quarter doesn't ensure that those things are meaningful. Listing objectives first, and then KRs based on those objectives, forces you to think about where you want to go before you think about how you want to get there.

At the end of each quarter, we grade each of our KRs from 0.0 to 1.0 depending on how well we met each goal. Setting measurable KRs is important because then you can grade them quickly and objectively, unlike a qualitative goal such as "Get better at [blah]" (what does "better" mean? How do you know when you've reached it?). If your key result was to write 5 blog posts and you only wrote 3, you score a 0.6. Easy. Average each set of KRs to see how well you met each objective, and average all your KRs to see how well you met your goals for the quarter.

But maybe the thing I like the best—especially when I start to think about OKRs outside of the workplace—is that you're supposed to set stretch goals. At Google, if you're performing well and meeting expectations, you should be hitting ~0.7's on your OKRs. Fully meeting an OKR (scoring a 1.0) means you went above and beyond and accomplished something extra special. This encourages people to set aggressive goals and to try to achieve more than they think is likely. It's actually looked down on if you get too many 1.0's too many times in a row—it means you're not setting aggressive enough goals for yourself.

I used to make New Year's resolutions in the half-assed way that many of us do: as an overly-optimistic wish list of what I'd like to become in the next year (more fit, more successful, less stressed), rather than a list of things I intended to achieve. I'd think big but inside I secretly wouldn't expect myself to be able to achieve that much, so I'd never seriously try to reach the goal. The problem with New Year's resolutions is that we see them as binary: in December, you can either say "I achieved that resolution this year," or "I did not achieve it." If you resolved to go to the gym 3 times/week and at some point you missed a week or two, you've already failed that resolution and there's less incentive to keep going for the rest of the year: even if you go 3 times/week for the rest of the year, you won't have fully met your resolution. This type of black-or-white assessment of goals practically guarantees that you're setting yourself up for disappointment.

For the last several years I've set something that's closer to "New Year's OKRs." They're stretch goals, and I think of "optimum performance" as scoring 0.7 on them. If I resolve to go to the gym 3 times/week and I miss a week or two, I'm still scoring a 0.96 on this goal, which is damn good! In fact, I could miss 13 weeks (hey, stuff happens) and still score 0.7 and feel good about myself. I've also noticed myself setting more and more measurable goals, such as "Design & publish 4 knitting patterns" rather than just "Design more patterns." It feels really good when you can score a solid 1.0 on a goal and know that it's not just because you're allowing yourself an overly-generous interpretation of "more" or "better." I still set qualitative goals too—be less judgmental, take the bus—but I feel like the score-things-as-OKRs-rather-than-resolutions perspective has helped me feel like I can set aggressive goals for myself and actually expect myself to follow through on them. Now, even if I feel guilty about slipping up on a resolution, that one slip isn't able to derail my ability to follow through on the resolution for the rest of the year and to feel good about my accomplishments at the end.

Here are some of my New Year's OKRs for 2011:

  • Finish knitting 3 sweaters
  • Design & publish 1 sweater pattern
  • Knit 11 shawls
  • Buy more often from independent yarn dyers
  • Get my unfinished knitting projects under control (I'd like to have ≤ 3 at the end of the year. Don't ask how many I have right now!)
  • Be less judgmental
  • Take the bus to work in the summer

What are yours?

Monday, December 6, 2010

Urban Craft Uprising recap

Every time I go to a street fair or art fair I get a bunch of great business cards from the merchants & crafts people, expecting to blog about it... and then never do. Today I'm rectifying that with a recap of my two (yes, two!) days spent at Urban Craft Uprising this weekend!

Well, maybe not a recap. More like a shopping list for the future. :-) I got a lot of great Christmas gifts from the vendors, but of course there's always so much more you want to buy than you can really justify. Happily, there were a ton of people at Urban Craft Uprising and it was really great to see so many supporting our local artists and craftsters!! It almost makes one optimistic about making a living in the handmade business.

I did most of my shopping on Saturday, but was so impressed with the vendors that I came back early on Sunday to wait in line for a swag bag. My two hours in the cold ended up being well worth it.

Here are some of the vendors that most caught my eye:

  • Imps & Monsters (Justin Hillgrove)
    Very cute & unique prints & paintings. A lot of his work has a dark or lonely quality to it, but then you'll find one that you can't help but smile at. I may or may not have gotten a few of these as Christmas presents. :-)
  • Texture
    Comfy, classic, casual-chic clothing made in Bellingham. After ooh-ing and aah-ing over the booth on Saturday, I went back on Sunday and got this irresistible skirt. Their palazzo pants looked pretty cozy, too.
  • Belle Epoch
    OMG feathers! This was another booth that I drooled over on Saturday and came back to on Sunday. They collect molted feathers from fowl and make them into rockin' jewelry and hair accessories. From long & dramatic to irresistibly iridescent, their stuff is eye catching in every form.
  • Queen Bee
    Everyone I know was drooling over her baby-soft faux leather bags and their beautiful embroidery. I was lucky enough to get this adorable coin purse in the swag bag I got on Sunday (squee!). When I went over to thank her I think she giggled at how excited I was. :-)
  • The Sprinkle Factory
    Jewelry that looks like CANDY! This stuff was seriously adorable and looked good enough to eat. They have all manner of rings, necklaces, clips, charms, etc. that look like cookies, cakes, lollipops, sushi, donuts, and more. I got my mom a cute necklace that I think her kids will like (she teaches at an elementary school).
  • Foamy Wader
    Delicate, sophisticated gemstone jewelry. I got a koi necklace here for someone who hopefully doesn't read this blog. :-P
  • Jewels Curnow
    I liked these folks' gemstone rings, and had an interesting chat with Robbie. This full lotus ring was one of my favorites. They also had these cool rings where the stone is set off to the side so it looks like it's sitting between your fingers.
  • Mermaid Empire (Rachel Rader)
    Rachel makes these great little flowers out of polymer clay—her earrings & brooches in particular grabbed me. I love her use of color—she stacks three or four gradations of color together to create a lot of depth in such small objects. I love combination of solid colors so this was right up my alley.
  • Bella Sisters
    Painfully adorable jackets. They take thrift store jackets & vests and make them into new, stylish, you'd-never-recognize-it fashion items by adding felt appliqués, lace, bustles, embroidery, and awesomeness. I didn't even look at the prices because I knew it would only make me sad... it was all lovely and oh-so-hip.
  • Slow Loris
    Screen printed clothing of all colors & shapes. This rockin' polo shirt dress caught my eye immediately. I'm not really into the bicycles, though, so I discovered that the base garment (minus the screen printing) is on sale for $17 right now...

Sooooo now you know what to get me for my birthday (you've got two weeks!). ;-)

Wednesday, November 3, 2010

Fondue

This pattern has moved to http://www.rogueedits.com/patterns/fondue-hat/.

Tuesday, December 22, 2009

Joey

This pattern has moved to http://www.rogueedits.com/patterns/joey-hat/.

Friday, August 7, 2009

Karo Socken

This pattern has moved to http://www.rogueedits.com/patterns/karo-socken/.

Saturday, June 6, 2009

S3 Performance Benchmarks

Over the last couple of weeks we've been working with S3 to read data to power real-time user query processing. So we've made a lot of optimizations and measurements of the kinds of performance you can expect from S3.

S3 Throughput: 20-40MB/sec (per client IP)
20 MB/sec is the neighborhood for many small objects, with as much as 40 MB/sec for larger objects. We're pushing a lot of parallel transfers and range queries on the same objects. Each request is only pushing about 200 KB/sec. I don't think I've ever seen a single connection push more than 5 or 6 MB/sec. I'm assuming this is partly S3 traffic shaping. So this should scales well if you've lots of clients.
S3 Response Time: 180 ms
We're pulling from EC2 (just across the hall from S3?) We've seen response time range between 8ms (just like a disk!) and as long as 7 or 8 seconds. But under 200ms is quite reasonable to expect on average. We're pushing a lot of parallel requests (thousands per second across our cluster, with hundreds on individual machines).
Parallel Connections to S3: 20-30 or 120
20-30 roughly maximizes throughput, 120 roughly maximizes response time. S3 seems to do some kind of traffic shaping, so you want to transfer data in parallel. If you're hosting web assets (e.g. images) at S3 this is less of an issue since your clients are widely distributed and will hit different data centers. But if you're serving complex client data requests pulling data from S3 at just a few servers, you might be able to structure your app to download data in parallel. Do it with 20-30 parallel requests. More than that and you start getting diminishing returns. We happen to run more than that (perhaps as many as 100 per process, with as many as 1000 per machine) because we're focusing on response time, rather than throughput.
S3 Retries: 1
We do see plenty of 500 or 503 errors from S3. If you haven't, just wait. We build retry logic into all our applications and typically see success with even just one or two retries with very short waits. I should recommend exponential back-off (that's what the Amazon techs say in the forums). So if you're making more than one or two retries, start waiting a second, then two, then four, and so on. I'd bail and send yourself an email if you don't get a 200 OK after four or five retries and a minute of waiting. But maybe retry the first one right away, it'll work 9.9 times out of ten :)

If you're getting different results, do let me know :)

Wednesday, March 25, 2009

Advertising fail

Amazon ad for hoity-toity

Not quite ironic enough for FailBlog, but still funny. Now if only the same ad appeared for [hanky-panky]...

Tuesday, March 17, 2009

Looking for a good home: Sammy

Sammy

This is Sammy. He's a foster cat from the Humane Society who's currently living with us to get a break from the shelter. He's totally adorable so I'm looking for a good home for him!

Sammy's been with us for nearly 2 weeks and he's been super low-maintenance. He's very friendly; he'll immediately walk up to new people rather than hiding or shying away. He has long whiskers and soft, silky fur which only gets silkier the more you pet him. He loves to be petted, especially on his face and belly, and he'll purr pretty much as soon as you touch him.

Sammy on my lap He seems to really like being around people, and is very amenable to whatever ways you want to give him attention. He doesn't mind sitting on your lap, being picked up, cuddled, used as a pillow, etc. He's also very happy to just sit in the same room as you while you're doing whatever; we've spent many hours sewing, knitting, and surfing the web together. He doesn't get freaked out by noises (like the sewing machine, or even the vacuum cleaner). I think he would make a great match for anyone looking for a lap cat, or a companion to just hang out with while watching movies and putzing around the house.

Incidentally, he doesn't do any of the bad stuff that my cats do, like chewing on cords or tipping over wastebaskets and playing in the trash. And he doesn't scratch the furniture (although he does like to knead his claws into the carpet when he's really happy).

If you or anyone you know is looking for a cat like this, please let me know (or contact the Humane Society directly)! You're also welcome to come over and meet him. More photos here.

Sunday, March 15, 2009

Performance Measurement for Small and Large Scale Deployments

As well as powering a few cool tools, Linkscape is a data platform. Performance (and its measurement) isn't just important to reduce user latency, or cut costs. It's actually something we're hoping is part of our core competency, something that adds significant value to our startup. And the shortest path to performance is measurement.

For those in a hurry, jump straight to the tools we're using.

This post is inspired by (and at times borrowed from) an email I sent to some friends for a consulting gig I did recently. But it rings so true, and I come back to it so often, that I thought I would share it. Alex and Nick, I hope you don't mind me sharing some of the work we've done on your very neat, very fun Facebook game.

Let me motivate the need for performance monitoring with a couple of case studies taken from our infrastructure:

This dashboard (above) illustrates 28 hours of load on our API cluster. I can immediately see service issues on the first server (the red segment of the first graph). This is correlated with a spike in CPU and some strange request patterns on the second server (the layered, multi-colored bar on the graph below). The degraded service lasted for a few hours, which was a configuration issue I solved in our monitoring framework. It should have guaranteed downtimes of no more than 4 minutes.

Even after solving our monitoring issue I still needed to investigate the underlying issue: I can see the CPU and request pattern are related. Ultimately I solved this issue within two weeks. Without this kind of measurement I would not even have known we had an issue, and would not have had the data to solve it.

From part of our back-end batch-mode processing, we had thought we'd tuned our system about as well as we could. At times we were pulling data through at a very respectable pace, roughly 10MB/sec per node. but we had also observed occasional unresponsiveness on nodes, with a corresponding slowness in processing. We left the system alone for a while, thinking, "don't fix it if it ain't broke". But recently we've been tuning performance for cost reasons; so we came back to this system.

Once we instrumented our machines with performance monitoring (illustrated above) we saw that the anecdotes were actually part of a worrying trend: the red circles show this. Our periods of 10MB/sec throughput are punctuated by periods of extremely high load. The graphs above show load averages of 10 or more on 4 core nodes, along with one process spiking up to hundreds of megabytes and nearly exhausting system memory. This high system load dramatically reduced our processing throughput.

It turned out that the load was caused by a single rogue program which consumed all available system memory due to buffered I/O. Usually we have a few I/O pipelines and give each many megabytes for buffering. However, this program has many dozens of pipelines, altogether consuming nearly a gigabyte of memory. This lead to significant paging and finally thrashing on disk.

Once we reduced the size of buffers (from roughly 40-100MB per pipeline to just 1-2MB per pipeline) we saw dramatic improvements in performance: a nearly 60% boost! And the nodes have become dramatically more responsive—no more load averages of 10+. The graphs above show load average maxing out at 4 and plenty of memory available. The data suggest that we might even be able to nearly double our performance with the same hardware by increasing parallelism and running another pipeline on each node.

All of this work is powered by simple monitoring and measurement techniques. Sometimes this has lead to significant, but necessary engineering work. But sometimes it's lead to a single afternoon's efforts yielding a 60% performance boost, with an opportunity to nearly double performance on top of that.

We're using a few tools:

  • collectd measures the system health dimensions (cpu, mem usage, disk usage, etc.) and sends those measurements to a central server for logging.
  • RRDTool records and visualizes the data in an industry standard way.
  • drraw gives me a very simple web interface to view and manage my visualizations.
  • Monit watches processes and system resources, bringing things back up if they crash and sending emails if things go wrong.

These tools work together, in an open, plug-in powered way. I could swap out individual components and move to other tools, such as Nagios (which I've used for other projects) or Cacti (which I have not used).

Whether you're an on-the-ground operations engineer looking to watch system health and fix issues before they turn into downtime, or you're managing large-scale engineering, looking to cut costs and squeeze out more page or API hits, these tools and techniques point you in the right direction and give you hard data to justify your efforts after the fact. We've had many high ROI efforts initiated and justified by this kind of measurement.

Sunday, March 8, 2009

Why is this Report So Slow: Let the Database Handle the Data

We have a data-rich report in our Linkscape tool with even more in our Advanced Report. We think the data is great. But the advanced report can be awfully slow to load. Don't get me wrong, we think it's worth the wait. But this kind of latency is a challenge for many products, and clearly, there's room for improvement. We're finding improvements by porting logic from the front-end into the data layer, and by paging through data in small chunks.

We present our data (links) in two forms. One is an aggregated view, showing the frequency of anchor text, one attribute of each link:

We also present a paged list of links, showing all the attributes we've got:

The time we spend on each request is very roughly illustrated by this diagram. From it you can see each component in our system: disk access, data processing, and a front-end scripting environment. I've included the aggregate time the user experiences as well. We have a custom data management system rather than using a SQL RDBMS such as MySQL. But I list it as SQL because SQL presents the same challenge.

In total the user can experience between 15 seconds to three minutes of latency! The slowness comes from a couple of design flaws. The first is that we're doing a lot of data processing outside our data processing system. Saying that programming environment doesn't matter is a growing trend, which has some advantages; rapid development comes to mind. But (and I'm a back-end data guy, so I'm a bit biased) it's important to let each part of your system do the work it's best at. For presentation and rapid prototyping that means your scripting environment. But for data that means data processing.

We're currently working on moving data processing into our data layer, resulting in performance something like that illustrated in the diagram below. The orange bars represent time spent in this new solution; the original blue bars are included for comparison.

In addition to latency improvements, pulling this logic out of our front-end adds that data to our platform and consequently makes it re-usable by many users and by applications. The maintenance of this feature will then lie in the hands of our data processing team, rather than our front-end developers. And we've taken substantial load off of our front-end servers, in exchange for a smaller amount of extra load on our data-processing layer. For us this is a win across the board.

The other problem we've got is that we're pulling up to 3000 records for every report, even though the user has a paged interface. And those 3000 records are generated from a join which is distributed across our data management platform, involving many machines, and potentially several megabytes of final data pulled from many gigabytes of source data.

The other big improvement we want to introduce is to implement paging at the data-access level. Since our users already get the data in a paged interface, this will have no negative effect on usability. And it'll make things substantially faster, as illustrated (again very roughly) below. The yellow bars illustrate the new solution's projected performance. Orange and blue bars are included for comparison.

The key challenge here is to build appropriate indexes for fast, paged, in-order retrieval of the data by many attributes. Without such indexes we would still have to pull all the data and sort it at run-time, which defeats the purpose.

In the solution we're currently working on we've addressed two issues. First, we've been processing data in the least appropriate segment of our system. Process data in your data management layer if possible. Second, we've been pulling much more data than we need to show to a user. Only pull as much data as you need to present to users; show small pages if you can. The challenges have been to port this logic from a rapid prototyping language like Ruby into a higher-performance language like C or stored procedures, and to build appropriate indexes for fast retrieval. But the advantages of this work are substantial, and are clearly worth it.

These issues are part of many systems out there and result in both end-user latency problems, as well as overall system scalability problems. Fixing those problems results in higher user satisfaction (and hopfully higher revenue), and reduces overall system costs.

By the way, we haven't released anything around these improvements in performance yet. If you want to keep up to date on Linkscape improvements watch the SEOmoz Blog or follow me on Twitter @gerner.

Wednesday, March 4, 2009

High Performance Computing at Amazon: A Cost Study

In building Linkscape we've had a lot of high performance computing optimization challenges. We've used Amazon Web Services (AWS) extensively and I can heartily recommend it for the functionality at which it excels. One area of optimization I often see neglected in these kinds of essays on HPC is cost optimization. What are the techniques you, as a technology leader, need to master to succeed on this playing field? Below I describe some of our experience with this aspect of optimization.

Of course, we've had to optimize the traditional performance front too. Our data set is many terabytes; we use plenty of traditional and proprietary compression techniques. Every day we turn over many hundreds of gigabytes of data, pushed across the network, pushed to disk, pushed into memory, and pushed back out again. In order to grab hundreds of terabytes of web data, we have to pull in hundreds of megabytes per second. Every user request to the Linkscape index hits several servers, and pages through tens or hundreds megabytes of data in well under a second for most requests. This is a quality, search-scale data source.

Our development began, as all things of this scale should, with several prototypes, the most serious of which started with the following alternative cost projections. You can imagine the scale of these pies makes these deicions very important.

These charts paint a fairly straight forward picture: the biggest slice up there, on the colocation chart, is "savings". We spent a lot of energy to produce these charts, and it was time well spent. We built at least two early prototypes using AWS. So at this point we had a fairly good idea, at a high-level of what our architecture would be, especially the key costs of our system. Unfortunately, after gathering quotes from colocation providers, it became clear that AWS, in aggregate, could not compete on a pure cost basis for the total system.

However, what these charts fail to capture is the overall cost of development and maintenance, and many "soft" features. The reason AWS was so helpful during our prototype process has turned out to be the same reason we continue to use it. AWS' flexibility of elastic computing, elastic storage, and a variety of other features are aimed (in my opinion) at making the development process as smooth as possible. And the cost-benefit of these features goes beyond the (many) dollars we send them every month.

When we update our index we bring up a new request processing cluster, install our data, and roll it into production seamlessly. When the roll-over is complete (which takes a couple of days), we terminate the old cluster, only paying for the extra computing for a day or so. We handle redundancy and scaling out the same way. Developing new features on such a large data set can be a challenge, but bringing up a development cluster of many machines becomes almost trivial on the development end, much to the chagrin of our COO and CFO.

These things are difficult to quantify. But these are critical features which make our project feasible at any scale. And they're the same features the most respected leaders in HPC are using.

All of this analysis forced us to develop a hybrid solution. Using this solution, we have been able to leverage the strength of co-location for some of our most expensive system components (network bandwidth and crawling), along with the strengths of utility computing with AWS. We've captured virtually all of the potential savings (illustrated below), while retaining most of our computing (and system value) at AWS.

I would encourage any tech leaders out there to consider their system set-up carefully. What are your large cost compentents? Which of them require the kind of flexibility of EC2 or EBS? Which ones need the added reliability of something like S3 with it's 4x redundancy? (which we can't find for less anywhere else). And which pieces need less of these things? Which you can install in a colocation facility, for less?

As an aside, in my opinion AWS is the only utility computing solution worth investigating for this kind of HPC. Their primitives (EC2, S3, EBS, etc.) are exactly what we need for development and for cost projections. Recently we had a spate of EC2 instance issues, and I was personally in touch with four Amazon reps to address my issue.

Saturday, November 1, 2008

Hhffrrrggh: WTF?

Once upon a time, while driving up to Madison from Chicago, I noticed a sign alongside the highway. It was one of those signs that precedes an exit and shows all the different restaurants that you'll find off that exit. I sped past at 80mph, barely glancing at the sign, and immediately thought Did I just see... ??. One of the restaurant adverts was pink with white text, so it was kind of hard to read, but I could've sworn it said something completely unintelligible. It was a weird bit of surreality in an otherwise uneventful drive, and I soon forgot about it.

But I drive that highway several times a year, and (when I remember) I started looking for that sign each time I drove past, trying to figure out whether I was crazy or whether it actually was complete gibberish. And today I'm pleased to report that I've found the culprit, and that I am not, in fact, crazy (or at least not hallucinating):

Welcome to the Hhffrrrggh Inn - Janesville's hmost hfun place to eat and drink. Can't say it, can't spell it, can't forget it.

So... did they just let their cat walk on the keyboard and name the restaurant after the results?

Saturday, October 25, 2008

Knitting breakthroughs

Ahh, the joys of being self-taught.

This week while poring over diagrams of a new knitting stitch I'm trying to learn, I realized that for the five years I've been knitting, I've been doing the most basic stitch—"the knit"—wrong.

You're doing it wrong.

Sigh. At least I figured it out before making a mess of my latest project. This is the first time I've tried anything that wasn't just a basic stockinette or seed stitch, so it never really mattered before.

Incidentally—although I don't think this had anything to do with my learning the stitch wrong—I taught myself to knit while I was in Paris, so my book is all in French. This means I don't really know any knitting vocabulary in English. For those interested, here's my new stitch:

My knitting project

The yarn is Plymouth baby alpaca grande paint, #8819. It's a bit expensive but is gorgeously soft (and the color is much better than this photo makes it out to be). Here's the stitch ("Grille ondulée"):

Cast on a number of stitches evenly divisible by 12.
1er rang: *4 mailles croisées à droite (glisser 2 m. en attente derrière le travail, tricoter les 2 m. suivantes à l'endroit, puis les 2 m. en attente), 4 m. endroit, 4 m. croisées à gauche (glisser 2 m. en attente devant le travail, tricoter les 2 m. suiv. à l'endroit puis les 2 m. en attente)*, répéter de * à *
2e et tous les rangs pairs suivants: à l'envers
3e et 7e rangs: à l'endroit
5e rang: *2 m. end., 4 m. croisées à gauche, 4 m. croisées à droite, 2 m. end.*, répéter de * à *
Répéter toujours ces 8 rangs.

Friday, October 24, 2008

Mimes as traffic cops

I know it's a blogging no-no to just republish stuff if you don't have original commentary to add, but I was so floored by learning about this yesterday that I just have to share it with you:

A mime in Bogotá

[Antanas Mockus, the former mayor of Bogotá, used] mimes to improve both traffic and citizens' behavior. Initially 20 professional mimes shadowed pedestrians who didn't follow crossing rules: A pedestrian running across the road would be tracked by a mime who mocked his every move. Mimes also poked fun at reckless drivers. The program was so popular that another 400 people were trained as mimes.

NPR story—which is even more compelling than the above article—here, starting at 34:00.

Thursday, October 16, 2008

Lessons Learned while Indexing the Web

If you know what I've been working on for the last nine months, you might (correctly) suspect that I've learned a few lessons about developing large-scale (highly scale-able) and complex software. Let me share some thoughts I've got about the subject.

But before I begin, I should point out some aspects that make this a special project. I can't speak for the UI, and while everyone on the engineering team worked on the project in the final months, the development team—especially in the early stages—was small. Ben Hendrickson and I headed up architecture and the software efforts. We were the "core" team for the back-end efforts. So this made some things a lot easier. I'll comment more about this later.

Be Bullish (but Realistic) in the Planning Stages

One thing that helped us tremendously was to be broad and optimistic in the early stages. Doing this gave us a large menu of features and directions for development to choose from as plans firmed up and difficulties arose. I know the adage, "Under-promise, over-deliver." And there's a place for that mentality. When we did start to firm up plans, clearly we were not going to promise everything. But we tried to keep things as fluid as possible for as long as possible. This worked out for almost all of our features, and by the time we were half-way through with our project we had the final feature set nailed down, and prototyped out.

There was one substantial feature that we had to cut just a few weeks before launch. Perhaps we were too bullish, but I believe that this kind of thing is fairly normal for large software projects. I'm looking forward to working on that for the next release. ;)

Have Many Milestones and Early Prototypes

I wish we had had more milestones and stuck to them. The last couple of months were hellish with literally 14-16 hour days, 7 days a week. With my commute, there were many days that I arrived home, got into bed, woke up, and rolled back onto the bus to start it all over again. Despite reading about "hard-core" entrepreneurs who have this as a "lifestyle", I would not recommend it for a successful software engineering team.

We hit our earliest milestones and even had an early version about four months before launch. But the two milestones between that prototype and launch both slipped and no one stepped in to repair the schedule. So the rest of the team (including Ben and me, plus another six software engineers) had to take up the slack at the last minute. Missing these milestones should have told us something about the remainder of the schedule. Frankly, I think we were lucky to launch when we did (good job team!).

Low Communication Overhead = Success

We were lucky to have a small team. For the back-end it was basically just Ben and me. And we do all of the data management and most of the processing in the back-end. So it was easy for Ben and me to stay in sync. Add to that the fact that we work well together and we were able to achieve in a small team, what normally requires a much larger team.

While I've worked in much larger organizations, I've never had the leadership role in those organizations that I do now. So I can't say how much of this advice applies to larger organizations. I guess my feeling is the same that many people have: keep related logic together in small teams, have clear interfaces to other units. This worked well when we integrated with our middle-ware and front-end.

KISS

Anna Patterson describes a simple roadmap for building search engine technology. I'm not saying we followed this plan, but I can say that our (I hope successful) plan is equally simple. Identify the work you need to do. Come up with reasonable solutions. Plan and implement them. Don't get bogged down in hype or fancy technology.

There's certainly more to success than these points. These are just what come to mind when I think about the success of this project. In any case, good luck on your own projects!

Wednesday, October 1, 2008

What's a little impending doom among friends?

As you hopefully know, CERN's Large Hadron Collider—one of the most elaborate physics experiments ever built—is finally finished and was first turned on several weeks ago. Leaving aside the fact that it broke down a few days later (!), I've been truly surprised by how many people are worried that it's going to create some sort of time-space anomaly that could destroy the world. The subject came up recently over lunch and (to my surprise) the majority of my PFM sisters were freaked out about our planet's impending Swiss-wrought doom.

And they're clearly not the only ones, since the BBC broadcast an interview the day before the collider was turned on in which they asked a scientist about the possibility of black holes being created during the experiments. His response was to laugh knowingly and say not to worry; even if the experiments do create (tiny) black holes, "there's no chance of them devouring the world. Ha, ha, ha!"

I couldn't find the exact interview, but here's a very similar one. Unfortunately it doesn't quite recreate the awkwardness of that scientist's particular response. I think the definition of 'egghead' somehow fundamentally involves the idea of believing so deeply in science that you'll laugh at another person's fear of death as if they were confusing a Stephen King novel for reality. "Black holes? Ha! Next thing I know, you'll be worried about langoliers!"

Monday, September 29, 2008

TSA Permitted & Prohibited Items

I'm flying to WI later this week (roller derby Eastern Regionals, baby!), and—having recently started a new knitting project—was wondering what happens when you try to bring knitting needles through airport security. Even though you could do much more damage with a ballpoint pen than with a blunt knitting needle, I would hate to underestimate TSA's overzealousness in "protecting public safety" in a post-9/11 world.

So I found this useful list of what's allowed and prohibited on airplanes. It even breaks things out into what's allowed in carry-ons vs. what's allowed in checked luggage. According to the list, knitting needles and crochet hooks are allowed on the plane; however, this follow-up article isn't exactly confidence-inspiring ("In case a Security Officer does not allow your knitting tools through security it is recommended that you carry a self addressed envelope so that you can mail your tools back to yourself as opposed to surrendering them at the security checkpoint").

[ Edited 11/24/2010: Just looked at the knitting/needle-crafting-specific article and it now says unequivocally that knitting needles and tools are allowed in all luggage! No more "we may or may not take them away from you." ]

I was surprised to learn that disposable razors and scissors < 4" long are allowed in carry-on luggage. Happily, the list confirms that throwing stars, cattle prods, hand grenades and tear gas are not.

I feel safer already.

[Edit: Maybe I just need one of these. "Nothing to see here, folks!" (Hat tip to Nish.)]

Saturday, September 27, 2008

Read a banned book

Today is the first day of Banned Books Week 2008. From their website:

Banned Books Week is the only national celebration of the freedom to read. It was launched in 1982 in response to a sudden surge in the number of challenges to books in schools, bookstores and libraries. More than a thousand books have been challenged since 1982. The challenges have occurred in every state and in hundreds of communities.

To my surprise, I discovered that a book I just started reading today (The Perks of Being a Wallflower, which I learned of through NPR) was one of the top 10 most challenged books in 2007. So I'll be celebrating Banned Books Week by curling up on the couch to finish it.

If you too value the freedom to access the literature of your choice—literature that may educate, entertain, shock, or open your mind—then check out this list of most frequently challenged books, visit your local library, and go exercise your First Amendment rights.

Friday, September 26, 2008

Ladies!

Thanks to beatnikside, I just stumbled across a delightful video which combines two very delightful things: roller skating and Flight of the Conchords! If you're unacquainted with either one, I highly recommend both. :-)