My team and I we are refactoring a REST-API and I have come to a question. For terms of brevity, let us assume that we have an SQL database with 4 tables: Teachers, Students, Courses and Classrooms. Right now all the relations between the items are represented in the REST-API through referencing the URL of the related item. For example for a course we could have the following { "id":"Course1", "teacher": "http://server.com/teacher1", ... } In addition, if ask a list of courses thought a call GET call to /courses , I get a list of references as shown below: { ... //pagination details "items": [ {"href": "http://server1.com/course1"}, {"href": "http://server1.com/course2"}... ] } All this is nice and clean but if I want a list of all the courses titles with the teachers' names and I have 2000 courses and 500 teachers I have to do the following: Approximately 2500 queries just to read the data. Implement the join between the teachers and courses Optimize with caching etc, so that I will do it as fast as possible. My problem is that this method creates a lot of network traffic with thousands of REST-API calls and that I have to re-implement the natural join that the database would do way more efficiently. Colleagues say that this is approach is the standard way of implementing a REST-API but then a relatively simple query becomes a big hassle. My question therefore is: 1. Is it wrong if we we nest the teacher information in the courses. 2. Should the listing of items e.g. GET /courses return a list of references or a list of items? Edit: After some research I would say the model I have in mind corresponds mainly to the one shown in jsonapi.org . Is this a good approach?

Reputation: 972

REST design principles: Referencing related objects vs Nesting objects

My team and I we are refactoring a REST-API and I have come to a question. For terms of brevity, let us assume that we have an SQL database with 4 tables: Teachers, Students, Courses and Classrooms. Right now all the relations between the items are represented in the REST-API through referencing the URL of the related item. For example for a course we could have the following

{ "id":"Course1", "teacher": "http://server.com/teacher1", ... }

In addition, if ask a list of courses thought a call GET call to /courses, I get a list of references as shown below:

{
   ... //pagination details
  "items": [
   {"href": "http://server1.com/course1"},
   {"href": "http://server1.com/course2"}...
 ]
}

All this is nice and clean but if I want a list of all the courses titles with the teachers' names and I have 2000 courses and 500 teachers I have to do the following:

Approximately 2500 queries just to read the data.
Implement the join between the teachers and courses
Optimize with caching etc, so that I will do it as fast as possible.

My problem is that this method creates a lot of network traffic with thousands of REST-API calls and that I have to re-implement the natural join that the database would do way more efficiently. Colleagues say that this is approach is the standard way of implementing a REST-API but then a relatively simple query becomes a big hassle.

My question therefore is: 1. Is it wrong if we we nest the teacher information in the courses. 2. Should the listing of items e.g. GET /courses return a list of references or a list of items?

Edit: After some research I would say the model I have in mind corresponds mainly to the one shown in jsonapi.org. Is this a good approach?

Upvotes: 1

Answers (3)

VoiceOfUnreason

Reputation: 57239

My problem is that this method creates a lot of network traffic with thousands of REST-API calls and that I have to re-implement the natural join that the database would do way more efficiently. Colleagues say that this is approach is the standard way of implementing a REST-API but then a relatively simple query becomes a big hassle.

Your colleagues have lost the plot.

Here's your heuristic - how would you support this use case on a web site?

You would probably do it by defining a new web page, that produces the report you need. You'd run the query, you the result set to generate a bunch of HTML, and ta-da! The client has the information that they need in a standardized representation.

A REST-API is the same thing, with more emphasis on machine readability. Create a new document, with a schema so that your clients can understand the semantics of the document you return to them, tell the clients how to find the target uri for the document, and voila.

Creating new resources to handle new use cases is the normal approach to REST.

Upvotes: 1

Roman Vottner

Reputation: 12839

My problem is that this method creates a lot of network traffic with thousands of REST-API calls and that I have to re-implement the natural join that the database would do way more efficiently. Colleagues say that this is approach is the standard way of implementing a REST-API but then a relatively simple query becomes a big hassle.

Have you actually measured the overhead generated by each request? If not, how do you know that the overhead will be too intense? From an object-oriented programmers perspective it may sound bad to perform each call on their own, your design, however, lacks one important asset which helped the Web to grew to its current size: caching.

Caching can occur on multiple levels. You can do it on the API level or the client might do something or an intermediary server might do it. Fielding even mad it a constraint of REST! So, if you want to comply to the REST architecture philosophy you should also support caching of responses. Caching helps to reduce the number of requests having to be calculated or even processed by a single server. With the help of stateless communication you might even introduce a multitude of servers that all perform calculations for billions of requests that act as one cohesive system to the client. An intermediary cache may further help to reduce the number of requests that actually reach the server significantly.

A URI as a whole (including any path, matrix or query parameters) is actually a key for a cache. Upon receiving a GET request, i.e., an application checks whether its current cache already contains a stored response for that URI and returns the stored response on behalf of the server directly to the client if the stored data is "fresh enough". If the stored data already exceeded the freshness threshold it will throw away the stored data and route the request to the next hop in line (might be the actual server, might be a further intermediary).

Spotting resources that are ideal for caching might not be easy at times, though the majority of data doesn't change that quickly to completely neglect caching at all. Thus, it should be, at least, of general interest to introduce caching, especially the more traffic your API produces.

While certain media-types such as HAL JSON, jsonapi, ... allow you to embed content gathered from related resources into the response, embedding content has some potential drawbacks such as:

Utilization of the cache might be low due to mixing data that changes quickly with data that is more static
Server might calculate data the client wont need
One server calculates the whole response

If related resources are only linked to instead of directly embedded, a client for sure has to fire off a further request to obtain that data, though it actually is more likely to get (partly) served by a cache which, as mentioned a couple times now throughout the post, reduces the workload on the server. Besides that, a positive side effect could be that you gain more insights into what the clients are actually interested in (if an intermediary cache is run by you i.e.).

Is it wrong if we we nest the teacher information in the courses.

It is not wrong, but it might not be ideal as explained above

Should the listing of items e.g. GET /courses return a list of references or a list of items?

It depends. There is no right or wrong.

As REST is just a generalization of the interaction model used in the Web, basically the same concepts apply to REST as well. Depending on the size of the "item" it might be beneficial to return a short summary of the items content and add a link to the item. Similar things are done in the Web as well. For a list of students enrolled in a course this might be the name and its matriculation number and the link further details of that student could be asked for accompanied by a link-relation name that give the actual link some semantical context which a client can use to decide whether invoking such URI makes sense or not.

Such link-relation names are either standardized by IANA, common approaches such as Dublin Core or schema.org or custom extensions as defined in RFC 8288 (Web Linking). For the above mentioned list of students enrolled in a course you could i.e. make use of the about relation name to hint a client that further information on the current item can be found by following the link. If you want to enable pagination the usage of first, next, prev and last can and probably should be used as well and so forth.

This is actually what HATEOAS is all about. Linking data together and giving them meaningful relation names to span a kind of semantic net between resources. By simply embedding things into a response such semantic graphs might be harder to build and maintain.

In the end it basically boils down to implementation choice whether you want to embed or reference resources. I hope, I could shed some light on the usefulness of caching and the benefits it could yield, especially on large-scale systems, as well as on the benefit of providing link-relation names for URIs, that enhance the semantical context of relations used within your API.

Upvotes: 0

Tarlog

Reputation: 10154

Yes, I totally think you should design something similar to jsonapi.org. As a rule of thumb, I would say "prefer a solution that requires less network calls". It's especially true if amount of network calls will be less by order of magnitude. Of course it doesn't eliminate the need to limit the request/response size if it becomes unreasonable.

Real life solutions must have a proper balance. Clean API is nice as long as it works.

So in your case I would so something like:

GET /courses?include=teachers

GET /courses?includeTeacher=true

GET /courses?includeTeacher=brief|full

In the last one the response can have only the teacher's id for brief and full teacher details for full.

Upvotes: 0

REST design principles: Referencing related objects vs Nesting objects

Answers (3)

Related Questions