Showing posts with label Requester. Show all posts
Showing posts with label Requester. Show all posts

Saturday, April 5, 2008

Requesters: Getting High Quality Results, Part 2

As a follow-up to my original post, Requesters: Getting High Quality Results, there are several more strategies Requesters can use to keep their quality high. You might consider these to be advanced strategies.

  1. Use groups of HITs to "grade" workers:
    Rather than using a short qualification test, Requesters can use HITs to grade workers. To do this, you can run several batches of HITs at various times of the day, and analyze the returned data carefully. Then, assign the top workers a qualification value, and only allow workers above that threshold to work on your future HIT groups. Amazon Media Content recently tried this method for their advanced HITs. The downside of this grading method is that you only catch those workers who happened to find your HITs when you posted them. In addition, you might not have a large enough group of workers to draw on, and over time you can have dwindling numbers of workers completing your HITs. To rectify these problems, you will occasionally have to post open HITs to re-grade workers.
  2. Use sliding grade scale:
    This is like using a dynamic qualification test. After setting a grade (either via a qualification test, or via an assigned value), Requesters can assign qualification points for every approved response, and take away points for every rejection. This way a worker is motivated to give the highest-quality responses. CastingWords transcription HITs work this way. Don't forget to tell workers this is your methodology!
  3. Include/exclude countries
    Besides qualifying workers based on their approval/rejection stats, Requesters can restrict or allow workers based on their location. Using a locale qualification is a blunt instrument, but sometimes it's the only instrument available.

    At least one Requester has previously mentioned that upon analysis, they noticed poor quality work returned from particular countries, perhaps due to language issues. Unfortunately, there is no Amazon "system qualification" for language. When a worker looks at their locale qualification, it states:
    The Location Qualification represents the location you specified with your mailing address. Some HITs may only be available to residents of particular countries, states, provinces or cities.
    Unfortunately, there is no "OR" operator when you list locale codes. This means that, as stated in the AWS mTurk Requester documentation, you cannot, for example, allow workers from US or GB or AU or NZ or CA. The best you can do is to exclude workers from a list of countries. What's unfortunate yet again, is that Amazon restricts the number of qualifications for a HIT group to a maximum of 10. (See this document and search for "QualificationRequirement".)

    Establishing HITs based on the location of the worker is fraught with political problems. Workers frequently complain on the message boards at Turker Nation when a Requester only allows workers from a particular country (especially the US). Using a locale qualification to effectively impose a language qualification will inevitably unfairly block some works who speak the language fluently. Be aware that some workers will be annoyed by this, but you might not have any other way to apply a fluency test.

As a Requester, you are well within your rights to include and exclude workers who don't meet your criteria to help you achieve the highest quality results. Nowhere in the terms of service does it state that you have to be "fair" in excluding workers. In fact, Amazon won't arbitrate in any Requester-Worker disputes, as stated in the Participation Agreement.

However, the more restrictions and complicated hoops you place in your process, the less likely a worker will work on those HITs. You should also be prepared to respond to many inquiries any time you exclude any group of people. If it appears that you aren't being fair, you could also get a bad reputation, which will scare off other workers.

Although you may have to test out different strategies, when implemented correctly these 11 strategies can help Requesters achieve a high quality return on investment. Workers who achieve the "elite" qualifications take pride in their status, and might well put your HIT groups at the top of their to-do list.


Wednesday, April 2, 2008

Requesters: Getting High Quality Results

Sounds like TagCow has had some great publicity and response to their enterprise! This means more work, but a few snags as well. Georgetag posted some questions to Turker Nation. These are problems that most Requesters will encounter, but particularly those with high-volume HITs. Although my response is particularly about the Image Tagging set, the tactics outlined below are general enough to be used by any Requester.

Quote:
Garbage tags (intentional
and unintentional):
We filter out meaningless words (the, it, an, a, with) but we are getting some totally irrelevant tags
Vulgarity (jokesters are hijacking the program)"
It looks like MTurk is doing some filtering and we are doing some filtering as well. (Can anyone confirm that?)
...
Incomplete taggings:
We have gotten some images tagged with "boy" where there is more that could obviously be said about the photo, like "boy", "playing", "trains"

There are several things you can do to keep the quality high. Here is a list of tactics implemented by other Requesters on mTurk. I'm certainly not suggesting you use all these ideas, but one or two might work well.
  1. Have a good, representative list of examples, and a good description.
    I know folks over at Turker Nation have already mentioned this, but your description is very vague. We aren't certain if you want us to make a list of everything we see in the picture, or just keep it as simple as possible. A separate webpage with lots of examples will go very far. We can then emulate these examples.

  2. Warn workers what response will be rejected, and what behavior will get them banned.
    A clear (but not overly-dramatic) warning might just be enough to scare off Roboform-type workers.

  3. Require a minimum approval rating for workers.
    Some recent HITs by Amazon's Media Content, and the information extraction group had the approval rating greater than 90%. (Smart Travel Media also has this threshold.) This sounds about right to me. After doing 40k HITs, mine is 99.8%. In the forums, even those who do HITs with high rejections (like the items HITs) still seem to have above 90%. This will prevent some workers who are continually trying to "game" the system.

  4. Offer a bonus for high-quality work.
    Georgetag is already offering up a volume bonus based on approvals, which is excellent! Few requesters do this, and more should. Once you get a good verification workflow established, you could track workers' responses and reward those that have submitted the highest quality and volume. Money talks.

  5. Include "gotchas."
    Include some pictures that should have an obvious response, like a flower, bird, etc., and particularly ones with a word to include. Then you can start to weed out or ban workers who miss these images. The "are these items different" has a qualification set up and a large set of gotchas. When you miss one, your qualification goes down by 200 points. Basically, after getting 3 wrong, you get timed out for some length of time.

  6. Ban very bad workers.
    And ban them quickly. If you haven't already implemented it, check if a single worker is giving you the same response (or few responses) over and over again. I wouldn't be surprised if workers are bypassing the "Enter ALL text found in image" step. Luckily this can be automated. And don't forget to give them some rejections if they replicate their responses beyond some acceptable level.

  7. Set up verification HITs.
    After getting all the tags for a given image, you could then create a HIT where the worker verifies that the tags are relevant, and could let you know if any are vulgar or meaningless. An image with a row of checkboxes would be ideal, where you select any tags that are bad, and a comment field to let you know about anything unusual. Hopefully then you can get a great set of tags and you can identify the bad workers using another method.

  8. Use a qualification test.
    This one might be a pain to grade, but you could have 5 images that the worker has to successfully tag before being able to do your HITs. Some qualifications even have a quiz about the purpose of the HIT and whether an example is appropriate or not. The quiz-style could be automatically graded.


Implementing 1, 2, and 3 is dead-easy. Making a nice page with a list of good and bad examples will go a long way to fixing some of your problems. I think some workers might be inadvertently giving you poor tags due to lack of understanding.

Politically, using any of the tactics 2-8 can be slightly tricky. (Item 1 should be done by all Requesters. Don't forget, the "description" field cannot be seen by workers once they are in the HIT.) You don't want to scare off your best workers, nor stop people from trying your HITs. Don't be too threatening, or too strict. Workers get very upset if they feel wrongly slighted, and will happily share with all on the Turker Nation forums.

To end on a positive note:
You'll find that most workers really do want to give you exactly the high-quality response that you require. When paid well, we are eager to perfect our responses, and love having discussions and feedback. Continue a good dialog, and you will have a group of willing, quality workers in no time!


Thursday, March 27, 2008

Resources Beyond mTurk

There are several places beyond mTurk that may help you in your Turking quests:

  1. Turker Nation: A mature, well-established forum for Workers.
  2. Mechanical Turk Sandbox: The official test & development site. Everything looks and acts as if it’s the regular mTurk, but you don’t get paid for doing or posting HITs. It’s like monopoly money. It will prompt you to create an account on the Sandbox if you want to play around. However, you just use your Amazon login yet again.
  3. Amazon Web Services Developer Connection -Mechanical Turk Forum: The official forum for mTurk Requesters. Very technical discussions about how to post HITs, the mTurk APIs, and other Requester coding issues. Amazon responds to technical questions here very quickly. Also, Amazon posts update notifications here. Useful for workers to get an understanding of what Requesters go through.
  4. Turkers Forum: A new forum created at the end of February 2007. Doesn’t have too many discussion threads yet, but it’s worth keeping an eye on. Appropriately enough, this Forum advertised and started populating threads via a HIT on mTurk.

Don’t forget as well, that many HITs on mTurk are through companies whose business models rely on mTurk. These Requesters might have more resources and information on their home pages.

If you’re thinking of posting your own HITs, these resources are specifically for the Requester side:

  1. AWS Developer Forum: (mentioned above)
  2. Command Line Tools: Open source toolkit to make writing HITs much easier. Download through svn.
  3. mTurk API Tools in Other Languages: PHP, Perl and Java.
  4. HIT-Builder.com: Hit-Builder offers more features than the standard mTurk requester interface. You can also hire them to help develop your HITs.
  5. von Kempelen: Company offers HIT designing and posting services.
  6. Dolores Labs: Offers HIT data collection services. Has many examples of test HITs.

If you know of any other useful Turking resources, post a comment, and if it’s worthy, we’ll add it to our list!


Wednesday, March 26, 2008

Glossary

Like all popular pastimes, mTurk has its own lingo. Here’s a little guide.



As new topics come up, I will add more terms to the glossary. If you have any to add, feel free to let me know!



To open hyperlinked terms in a new tab, hold down the CTRL key and left-click the name.)



URL for Google Spreadsheet: http://spreadsheets.google.com/pub?key=pxpTLP5OADPhuuDGV0UDiFg



Monday, March 24, 2008

The six categories of HITs

There are numerous Requesters on mTurk. However, only a few types of HITs appear over and over again.

I have categorized HITs into 6 categories:

  1. Decision: In these HITs, the worker is asked to make a judgment call or to take a poll. Decision HITs usually contain radio buttons or a check list and are quick to perform. Examples of decision HITs are: Amazon's Are these items different, Powerset's Evaluate Search Results, and Content Review's Review User Submitted Images.

  2. Research: For research HITs, the worker must search for information, generally on the web. The response is typically a URL, copied reference material, or data. Examples of research HITs are Amazon's NowNow Research Questions, Unspun's Find a URL/Amazon product identifier, and ClayValet's Find a product group of HITs.

  3. Image Tagging: Workers interact with the picture in some way. Usually the worker is marking a set of specific features in the image. The two most common Requesters in this sub-type are Geospatial Vision (marking road features) and True Yardage (marking features on golf courses).

  4. Transcription: Workers are asked to transcribe text from an audio or video file. HITs that ask workers to transcribe text from an image are also included.

  5. Create: Generally a more involved HIT, these HITs require the worker to create some original content. There are many types of creation HITs. Some ask you to write an article or rewrite a sentence (e.g. ContentSpooling, Paul Pullen), create trivia questions (e.g. UQsoft), or draw something (e.g. draw).

  6. Traffic Generator: These HITs usually are trying to generate traffic to their website. They might ask the worker to click through some links. Sometimes they require the worker to comment on a blog. Other traffic generator HITs want workers to post links back to their own websites.

These six categories of HITs encompass just about every HIT you encounter on mTurk. I'm sure there are some oddballs out there as well that don't fit these five. I just can't think of any!

You might argue that a few HITs straddle more than one category. For example, rewriting sentences generally means the Requester is using workers for Search Engine Optimization (such as ContentSpooling). Although this has the end result of traffic generation, the actual work performed is mostly a creative process for the worker. Likewise, you could argue that asking a worker to generate a list of tags for an image is a creative process -- as is crafting a response for NowNow questions. However, the former is mostly an image-based process with minimal effort and the latter takes tremendous amounts of research.

In the list of Requesters and HITs, the categories are based on the major type of work a Turker is asked to perform.

Sunday, March 23, 2008

Requesters

Here is a list of frequent Requesters, the types of HITs they offer, and roughly how quickly they pay when your HIT is accepted. I will update this list from time-to-time to reflect the current Requesters at mTurk.



note: This is not meant to be a comprehensive list! But it’s a good place to start.





URL for Google Spreadsheet:
http://spreadsheets.google.com/pub?key=pxpTLP5OADPhLyybn0fY_Ag



If you click on a Requester name, you will be taken (within the iframe) to a search page for all HITs currently listed by that Requester ID. To open a new tab with that search, hold down the CTRL button before clicking. Not all Requesters will have HITs listed.



Any cell with "??" indicates missing data. If you have any of the missing information, contact me and I will update the spreadsheet.



The categories listed in the "Type" column are based on the 6 categories of HITs, described here. In addition, the time to payment is an estimate based on my experience. Beware that Requesters may change their time to payment.


Subscribe to: Posts